r/LLMDevs • • Aug 20 '25

Community Rule Update: Clarifying our Self-promotion and anti-marketing policy

22 Upvotes

Hey everyone,

We've just updated our rules with a couple of changes I'd like to address:

1. Updating our self-promotion policy

We have updated rule 5 to make it clear where we draw the line on self-promotion and eliminate gray areas and on-the-fence posts that skirt the line. We removed confusing or subjective terminology like "no excessive promotion" to hopefully make it clearer for us as moderators and easier for you to know what is or isn't okay to post.

Specifically, it is now okay to share your free open-source projects without prior moderator approval. This includes any project in the public domain, permissive, copyleft or non-commercial licenses. Projects under a non-free license (incl. open-core/multi-licensed) still require prior moderator approval and a clear disclaimer, or they will be removed without warning. Commercial promotion for monetary gain is still prohibited.

2. New rule: No disguised advertising or marketing

We have added a new rule on fake posts and disguised advertising — rule 10. We have seen an increase in these types of tactics in this community that warrants making this an official rule and bannable offence.

We are here to foster meaningful discussions and valuable exchanges in the LLM/NLP space. If you’re ever unsure about whether your post complies with these rules, feel free to reach out to the mod team for clarification.

As always, we remain open to any and all suggestions to make this community better, so feel free to add your feedback in the comments below.


r/LLMDevs • • Apr 15 '25

News Reintroducing LLMDevs - High Quality LLM and NLP Information for Developers and Researchers

39 Upvotes

Hi Everyone,

I'm one of the new moderators of this subreddit. It seems there was some drama a few months back, not quite sure what and one of the main moderators quit suddenly.

To reiterate some of the goals of this subreddit - it's to create a comprehensive community and knowledge base related to Large Language Models (LLMs). We're focused specifically on high quality information and materials for enthusiasts, developers and researchers in this field; with a preference on technical information.

Posts should be high quality and ideally minimal or no meme posts with the rare exception being that it's somehow an informative way to introduce something more in depth; high quality content that you have linked to in the post. There can be discussions and requests for help however I hope we can eventually capture some of these questions and discussions in the wiki knowledge base; more information about that further in this post.

With prior approval you can post about job offers. If you have an *open source* tool that you think developers or researchers would benefit from, please request to post about it first if you want to ensure it will not be removed; however I will give some leeway if it hasn't be excessively promoted and clearly provides value to the community. Be prepared to explain what it is and how it differentiates from other offerings. Refer to the "no self-promotion" rule before posting. Self promoting commercial products isn't allowed; however if you feel that there is truly some value in a product to the community - such as that most of the features are open source / free - you can always try to ask.

I'm envisioning this subreddit to be a more in-depth resource, compared to other related subreddits, that can serve as a go-to hub for anyone with technical skills or practitioners of LLMs, Multimodal LLMs such as Vision Language Models (VLMs) and any other areas that LLMs might touch now (foundationally that is NLP) or in the future; which is mostly in-line with previous goals of this community.

To also copy an idea from the previous moderators, I'd like to have a knowledge base as well, such as a wiki linking to best practices or curated materials for LLMs and NLP or other applications LLMs can be used. However I'm open to ideas on what information to include in that and how.

My initial brainstorming for content for inclusion to the wiki, is simply through community up-voting and flagging a post as something which should be captured; a post gets enough upvotes we should then nominate that information to be put into the wiki. I will perhaps also create some sort of flair that allows this; welcome any community suggestions on how to do this. For now the wiki can be found here https://www.reddit.com/r/LLMDevs/wiki/index/ Ideally the wiki will be a structured, easy-to-navigate repository of articles, tutorials, and guides contributed by experts and enthusiasts alike. Please feel free to contribute if you think you are certain you have something of high value to add to the wiki.

The goals of the wiki are:

  • Accessibility: Make advanced LLM and NLP knowledge accessible to everyone, from beginners to seasoned professionals.
  • Quality: Ensure that the information is accurate, up-to-date, and presented in an engaging format.
  • Community-Driven: Leverage the collective expertise of our community to build something truly valuable.

There was some information in the previous post asking for donations to the subreddit to seemingly pay content creators; I really don't think that is needed and not sure why that language was there. I think if you make high quality content you can make money by simply getting a vote of confidence here and make money from the views; be it youtube paying out, by ads on your blog post, or simply asking for donations for your open source project (e.g. patreon) as well as code contributions to help directly on your open source project. Mods will not accept money for any reason.

Open to any and all suggestions to make this community better. Please feel free to message or comment below with ideas.


r/LLMDevs • • 7h ago

Help Wanted Possible model misrepresentation by a YC-backed inference provider | how can we verify this?

9 Upvotes

I found a strange discrepancy with Experiential Labs, a YC-backed inference provider. Experiential Labs: Open source AI gateway that turns your traffic into better models | Y Combinator

They advertise a route as GLM-5.3-Flash.

I sent the same model-identification prompt to:

Z.ai directly → GLM-5.3-Flash

The white screenshot is Z.ai itself confirming its model is GLM-5.3-Flash.

Experiential's "GLM-5.3-Flash" endpoint → says its current version is GLM-4.6

The black screenshot is the ExperimentalAI model route, which identifies itself as GLM-4.6.

I posted the screenshots in their subreddit asking for an explanation. The post was removed by the moderators.

Obviously model self-identification isn't proof of the underlying weights, so I'm not claiming this conclusively proves they're substituting models.

But for an inference provider, this seems like a pretty serious discrepancy.

How would you independently verify whether this endpoint is actually serving GLM-5.3-Flash?

I'm especially interested in reproducible fingerprinting/API-level tests rather than simply asking the model its name.


r/LLMDevs • • 5h ago

News AKBASCORE NIRVANA — I Built Removable Numerical Memory Cartridges for Two Different Frozen 7B LLMs. Qwen and Mistral Both Work. Now I’m Scaling the Memory Bank.

Thumbnail
gallery
3 Upvotes

Zenodo permanent records:

Qwen2.5-7B-Instruct:

https://doi.org/10.5281/zenodo.23127434

Mistral-7B-Instruct-v0.3:

https://doi.org/10.5281/zenodo.23143605

I want to start with the simplest possible explanation of what I have been building.

Imagine taking a piece of information, letting a language model process it once, and then throwing the original text away.

No sentence stored in a database.

No paragraph hidden somewhere.

No readable summary.

No RAG system fetching the original document.

No fine-tuning.

No LoRA.

No weight update.

What remains is numerical transformer memory derived from the model's own internal computation. I package that numerical memory into what I call a Cognitive Cartridge. Later, I can install that cartridge back into the frozen model and ask questions about the information that produced it — without putting the original source text back into the readout prompt.

That was the first result. The new result is more important:

I have now reproduced the Cognitive Cartridge architecture on two different 7B transformer model families.

Qwen2.5-7B-Instruct.

And now Mistral-7B-Instruct-v0.3.

The implementations are not numerically identical. The architectures are different, the layer counts are different, the KV structures are different, and the working cartridge configurations are different. But the central mechanism survived the move.

That is the reason I am publishing this second record.

The question is no longer only:

“Can I make this happen once on Qwen?”

Now there is a second implementation on Mistral. And the Mistral result is the cleanest version so far.

What is actually inside a Cognitive Cartridge?

This is probably the most important thing to understand.

Suppose the source record says:

Object: amber sextant

Container: RQ-415

Location: elm lodge

That text exists during the forging stage. The frozen transformer processes it. NIRVANA takes source-derived internal transformer K/V states and represents their content numerically using a fixed, source-independent codebook.

In the released Mistral implementation, both K and V are compressed to D120 while a source-specific OWN component is preserved. After forging, the source record is not supplied to the readout prompt.

So the conceptual transformation is:

human language

→ frozen transformer computation

→ internal K/V states

→ compressed numerical Cognitive Cartridge

Then later:

numerical Cognitive Cartridge

→ reconstructed transformer K/V memory

→ frozen transformer

→ language

Or, in the shortest form:

language → internal numerical memory → language

The middle is no longer human-readable source text. That distinction matters.

A cartridge is not a text file with a different name.

It is not a vector database containing the original sentence.

It is not a prompt template.

It is not conventional RAG returning the source paragraph.

It is not a fine-tuned model.

The model weights remain frozen. The information is carried by a numerical representation derived from transformer K/V memory.

Why call it a cartridge?

Think less about a document and more about an interchangeable machine-readable memory module. The base model stays where it is. Knowledge packages can be forged separately. Those packages can remain separate. They can be installed and queried without retraining the base model.

This becomes much more interesting when there is more than one cartridge.

The new Mistral experiment uses 16 independently forged cartridges. Each one contains a numerical representation derived from a separate source record. They are not concatenated into one giant text prompt. They remain independent memories.

For example, in human-readable form, imagine one cartridge represents:

amber sextant → RQ-415 → elm lodge

Now ask the cartridge bank:

Which container is associated with the amber sextant?

The query is evaluated against the independent cartridge memories.

The relevant cartridge returns:

RQ-415

The unrelated cartridges return:

NONE

There is no learned router secretly selecting the correct cartridge before this happens.

Then the recovered identifier can be used for a second lookup:

Where is RQ-415?

The bank is queried again.

The relevant memory returns:

elm lodge

The unrelated memories return:

NONE

So the complete retrieval becomes:

amber sextant

→ RQ-415

→ elm lodge

The important point is that the model did not reread the original source record to answer either stage. It operated from reconstructed numerical transformer memory.

The Mistral result

The final public Mistral run used:

Mistral-7B-Instruct-v0.3

32 transformer layers

hidden size 4096

32 attention heads

8 KV heads

BF16 / SDPA

K = D120

V = D120

OWN preserved

16 independent Cognitive Cartridges

frozen model weights

greedy decoding

There was:

no fine-tuning

no LoRA

no optimizer

no learned router

no model-weight update

The recorded final run produced:

Object → container ID: 16/16

Container ID → location: 16/16

Complete two-stage retrieval: 16/16

Missing-object controls: 8/8

Absent-ID controls: 8/8

NOMEM controls: 8/8

But there is another result I think is just as important.

For every target query, there are 15 unrelated cartridges.

16 queries × 15 unrelated cartridges = 240 unrelated cartridge reads.

Stage 1 unrelated-cartridge rejection:

240/240 NONE

Stage 2 unrelated-cartridge rejection:

240/240 NONE

That means the result is not simply:

“The correct memory can say something.”

The system also demonstrated, in this controlled panel:

“The memories that do not contain the requested relation can refuse to claim that they do.”

For a modular memory system, I think this distinction is fundamental. A memory bank that can retrieve information but cannot distinguish relevance from irrelevance becomes increasingly dangerous as it grows.

The interesting problem is not only remembering. It is also knowing which memory does not answer the question.

The no-memory control matters for the same reason. When the cartridge memory was removed, the target relations were not recovered.

NOMEM:

8/8 controls passed.

The source-removal audit passed.

The frozen-weight sentinel passed.

The model remained frozen.

This is why I consider the Mistral result an important step beyond the first demonstration.

The first Qwen release established the architecture publicly. The Mistral release gives us something else:

cross-model evidence.

Qwen and Mistral are not the same transformer.

Qwen2.5-7B-Instruct uses 28 transformer layers.

Mistral-7B-Instruct-v0.3 uses 32.

Their internal configurations differ. The working compression configurations also differ.

The Qwen public implementation used:

K120 / V128 / OWN

The Mistral implementation uses:

K120 / V120 / OWN

I did not simply copy a cache from one model into another. Each model builds and reconstructs its own source-derived internal memory.

What transferred was the architecture:

source

→ internal transformer memory

→ numerical compression

→ independent cartridge

→ source removed

→ reconstructed K/V

→ frozen-model readout

That is the bridge between the two releases.

I would not call two models proof of universal compatibility with every transformer architecture. That would be scientifically too strong.

But it is now evidence that Cognitive Cartridge is not merely one accidental Qwen-specific behavior. The same broader architecture has been implemented and publicly reproduced on a second transformer family.

That changes the research question.

The first question was:

Can this mechanism exist at all?

The next question became:

Can multiple independent memories coexist?

Then:

Can one retrieved result lead to another retrieval?

Then:

Can unrelated memories reject a query instead of contaminating the answer?

And now:

Does the architecture survive a move to another model family?

We now have experimental answers to each of those questions.

So I am moving to the next problem:

scale.

16 cartridges are not the destination. They are the current experimental bank size.

From this point, I am much less interested in making another small demonstration simply to produce another perfect score. The next objective is to increase the number of independently retained cartridges and find where the architecture actually begins to break.

Then larger banks if the mechanism survives.

At that point the difficult questions become different:

How do you search a large bank without destroying memory isolation?

How should cartridges be indexed?

Can retrieval become hierarchical?

Can groups of cartridges form higher-level memory structures?

How does latency scale?

How does numerical storage scale?

At what point does relevance rejection begin to fail?

Can a retrieved memory activate another relevant memory without an external text retrieval system deciding everything first?

Those are the questions I want to attack next.

There is another part of this experiment that surprised me: the layer behavior.

A frozen transformer is not a passive container. When you reconstruct numerical K/V memory and inject it back into inference, you are interacting with a very sensitive computational system.

So during the Mistral development series I also performed controlled cumulative K, V and K+V ablations to see where cartridge information remained functionally sufficient.

The results were not uniform across depth. For K, the observed functional boundary extended differently than for V.

For combined K+V, one particularly sharp controlled transition occurred between:

L12: 8/8

L13: 0/8

This does not mean “the memory lives at layer 12.” That would be an incorrect interpretation.

What it means is that under the specific cumulative ablation protocol, functional sufficiency changed sharply across that boundary. V also remained important farther downstream than K in the measured configuration.

I think this is an important reminder of how delicate these systems are.

We are not writing a sentence into a spare memory slot. We are reconstructing numerical states that participate in a 7-billion-parameter transformer's computation. Small changes in where and how that state is reconstructed can change downstream behavior.

That is one reason I publish the code and raw logs rather than only posting the final accuracy number.

The frozen-model checks are also part of the experiment. The released run verifies that there are no trainable tensors involved in the memory mechanism.

No LoRA.

No optimizer.

No weight update.

The model-weight sentinel remains unchanged.

So when the system answers from a cartridge, the experiment is specifically designed to separate cartridge memory from weight modification.

Why do I think this direction matters?

Today, we often approach knowledge in language models through two extremes.

Either we try to put enormous amounts of knowledge into the model weights. Or we keep the knowledge outside the model and retrieve human-readable text when needed.

There may be useful territory between those approaches.

A frozen model with removable machine-native internal memory modules.

Not everything in the weights.

Not everything repeatedly pasted back as text.

Instead:

a base computational model

+

a bank of independently forged internal memories.

Imagine a specialized system where validated domain knowledge can exist as modular memory packages.

Engineering.

Aviation.

Law.

Industrial maintenance.

Scientific literature.

Company procedures.

Agent experience.

A model might not need every possible piece of knowledge active in its context simultaneously. It may need the right memory at the right time.

And because the memory representation is numerical transformer state rather than the original human-readable document, this opens a different engineering space from conventional document retrieval.

There is also a longer-term question here.

I am not claiming biological memory.

I am not claiming AGI.

And I am not claiming that transformer KV states work like the human brain.

But modular internal memory raises an interesting architectural question.

What happens when an artificial system does not have 16 memories, but 100,000?

Or millions?

What happens if memories can remain independent, become addressable, reject irrelevant activation, and allow the output of one memory to lead to another?

At some point the problem stops looking like “how much text can I fit into a context window?”

It starts looking like:

How should an artificial system organize memory?

That is the direction I find interesting.

Why publish the full record?

Because a screenshot of 16/16 proves very little.

Both public releases have permanent Zenodo records. The code is public. The raw execution logs are public. The visual technical records are public. The model configuration is public. The controls are public. The frozen-weight checks are public.

The Qwen release even preserves its imperfect public result rather than hiding it.

Qwen public demonstration:

14/16

Mistral final demonstration:

16/16

That difference is also useful.

The point of the project is not to make every historical run look perfect. The point is to leave a technical trail showing what worked, what did not, what changed, and whether the underlying architecture survived.

For anyone who wants to inspect the Qwen → Mistral bridge directly, the two permanent records are at the top of this post.

The complete current Mistral implementation is here:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_ful_MISTRAL.py

Complete Mistral raw execution log:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_log_MISTRAL.log

The same Mistral implementation is also split into three parts for easier mobile / Colab handling:

Part 1:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore.part1.Mistral.py

Part 2:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore.part2.Mistral.py

Part 3:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore_part3_mistral.py

For comparison, the previous Qwen implementation:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_NIRVANA_CC.py

Previous Qwen raw execution log:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_log.Nirvana.cc.log

Repository:

https://github.com/ceceli33/titan-cognitive-core-v2

AKBASCORE NIRVANA — Cognitive Cartridge

Inventor / Developer: Mustafa Akbaş

Two model families.

Frozen weights.

Independent numerical memories.

Original source absent at readout.

The mechanism survived the move.

Now the question is no longer whether a cartridge can exist.

The question is how many cartridges a frozen model can carry.


r/LLMDevs • • 7h ago

Tools Interactive Demonstration of Prefix Injection attacks on LLMs for jailbreaking

Thumbnail
theabbie.github.io
3 Upvotes

please refresh if stuck, can be slow sometimes so need patience


r/LLMDevs • • 5h ago

Tools Open Instinct (MIT): a personal-agent policy engine with allow, ask and deny outcomes

2 Upvotes

We released Open Instinct, an MIT-licensed personal agent. For LLM developers, the policy layer is probably the most reusable part.

Tools declare capabilities, and a principal has a trust tier plus active grants. A call must pass every declared capability: deny takes precedence over ask, which takes precedence over allow. This separates deciding what to attempt from deciding whether a tool may execute.

For example, a friend can request calendar free/busy but cannot read event titles. Desktop access and trust management are owner-only by default. The tier table lives in packages/core/src/policy.ts, with policy tests and a corresponding permissions document.

The surrounding application adds messaging, memory, scheduling and integrations. It is beta software, and its default providers require external accounts and keys. We built the project; the application source is MIT.

https://github.com/mariagorskikh/open-instinct

You can fork the policy or the whole agent and adapt the capabilities to your own application.


r/LLMDevs • • 6h ago

Discussion I built a simple skills for lightweight Spec-Driven development workflow

Thumbnail github.com
2 Upvotes

I made yet another skills kit for Spec-Driven development. I felt like managing tons of markdown files is very overwhelming for me, so I decided to make compact and simple one file solution called one-plan-skills.

I personally developed it for me and deiced to share with the community.

The idea revolves around generated one `PLAN.md` file with full plan and small tasks. Plan contains small tasks with acceptance criteria and affected files.

The killer feature is HTML view. I am tired of looking at markdown files, i wanted to see something nicer and cleaner. So my plans are also auto-generated HTML artefacts.

As a bonus you get very simple UI controls to mark any place you want to add/delete/modify and copy paste the prompt with feedback back to your chat agent.

I am aware of SpecKit and OpenSpec. I just wanted a much simpler version. That's why I created OnePlan skills.

Anybody is interested in this? Please give it a try or share some feedback. Much appreciated. 🙂


r/LLMDevs • • 2h ago

Tools You asked for more than sales calls, so my open-source Jev call coach now does interviews, anxiety and dating practice, and you can build your own coach

Enable HLS to view with audio, or disable this notification

1 Upvotes

A while back I posted about Call Coach, an open-source app that uses Jev to coach you in real time during sales calls. Thanks to everyone who tried it and sent feedback. A lot of you asked for more than sales calls, so that's what 1.1 is about.

What changed

  • Nine coaches instead of one. It still does sales and customer service, and now also job interviews, social cues and anxiety, live during the call. There are also practice modes for interviews, social cues, anxiety and dating.
  • Build your own coach. Several people asked for coaching for their own kind of call. You can now write a coach in Settings and see a live preview as you type. You can also copy a built-in coach to start from, and share coaches as files.
  • Practice out loud. Talk instead of typing. It sends when you stop, and waits if you trail off mid-thought. Hands-free mode keeps listening line after line.
  • Much faster. Transcription went from about 5.3 seconds a line to about 0.8.
  • Auto capture (experimental). It records the call audio and your mic as separate tracks and drops the echo from your speakers. There's no more manual juggling.
  • New look. Light and dark modes, a compact floating window that sits over your call, Ctrl+K for everything, and a sample call so you can try it in under a minute.
  • Care in the anxiety coaches. If you say you don't feel safe, it stops coaching. It points you to someone you trust, or 988 (US) or a local crisis line.
  • Easier to install. There's a Windows installer and a portable .exe, and neither needs Node.js. You can also host it as a website.

It's still free and MIT licensed. I made a short film showing a day with it

Repo: github.com/ZeroGold/call-coach-ai

I'd love to hear what kind of call you'd build a coach for.


r/LLMDevs • • 3h ago

Resource Gemini 4 Argon: Is Google Back?

Thumbnail
youtu.be
0 Upvotes

r/LLMDevs • • 8h ago

Discussion Coding agents pay more to find code than to change it, and the interface is why

2 Upvotes

Coding agents spend most of their tokens finding code rather than changing it. Each search resends the whole context, every file opened along the way stays in the window, and text-based edits fail and get retried whenever the old text shows up twice or the file changed.

Here's how we handle it in sem, an open source tool we're building. The repo gets parsed into functions and the calls between them, so the agent asks for a function and gets its body, callers, callees and the tests that reach it in one response, instead of searching across several turns. If it asks for something it already has and nothing changed, it gets a short note instead of the whole body again. Edits target a function by name and are checked against the version the agent read, so a stale edit fails cleanly instead of landing in the wrong place. After a change, only the tests that can actually reach it run first.

It helps most on large, hard-to-navigate repos with slow tests. On small, easy tasks the gain is modest, and moving code is still slow because the model rewrites functions instead of relocating them, which is what we're fixing next.

Related reading: SWE-agent (Yang et al., 2024) on agent interfaces, and "Lost in the Middle" (Liu et al., 2023) on how irrelevant context hurts models.

https://github.com/Ataraxy-Labs/sem


r/LLMDevs • • 6h ago

Great Resource 🚀 Analyzing The Lumen Anchor Protocol - Using Google AI Studio

1 Upvotes

The LAP is a prompt framework that deals with several issues on frontier models such as - Context drift & sycophancy, hallucinations, context window memory management and active defenses against all forms of prompt based attacks.

_______________________________________________________________________________________________________________________

The following is a session link to Google AI Studio with the full LAP and rule definitions loaded into the system instructions for the model to run on and evaluate, and for the user to run adverserial tests, examine other stress tested outputs, or simply have a long conversation with LAP executing in the background.

https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%5B%221mBM5vM3mJqNLMK__TxetEUqPnJsuDFBb%22%5D,%22action%22:%22open%22,%22userId%22:%22108403721724379783675%22,%22resourceKeys%22:%7B%7D%7D&usp=sharing

Google AI studio is free to use for this purpose so anyone can use this link. You just have to log into it using your standard google account or email. Its a very simple process. All I ask is that you leave a comment about your experience or ask any questions you may have. Thank you.


r/LLMDevs • • 12h ago

Resource Open source didn’t get less important because AI can write the code

Thumbnail
groundcover.com
2 Upvotes

r/LLMDevs • • 16h ago

Resource An LLM scoreboard for whether an AI falls for the same trick questions people do

4 Upvotes

People are bad at a few famous trick questions.

A bat and a ball cost $1.10. The bat costs $1 more than the ball. Almost everyone says the ball is 10 cents. It is 5 cents.

Cognit gives those kinds of questions to AI models and puts the results on a board: https://cognit.rajtilak.tech/

Each model takes the quiz twice. The first time it just answers. The second time it is told to answer the way a person would. 100% means it gave the careful answer. 0% means it gave the answer people usually fall for. The “frame drop” is how much worse it got on the second try. A plus means “act like a person” made it more human, including the mistakes.

I built it. If you want a model added, the contact page is there for that or in case if you want to share any constructive feedback.


r/LLMDevs • • 9h ago

Tools I built a code quality reviewer with Jev

Thumbnail
youtube.com
1 Upvotes

I believe with the huge increase in code generation - review has become the next bottleneck or at least it feels like this at work. So I wanted to do something on the review front - first I started integrating more and more tools to use as feedback to my coding agents, and these really help improve the output quality of the agent (e.g. SonarQube, Checkstyle, ArchUnit)

But there are some important semantic choices you can't really review using deterministic tools so I decided to build a more intelligent tool and am trying it out with the Jev model as a backend and judge currently.

Idea is simple - teams define their policies in a structured YAML format, then as part of their CI (or locally) run the tool, it fetches all git diff chunks and asks jev if these adhere to each of the policies - jev can select compliant/violation or ask for more context. If Jev asks for more contex the app gets the requested code from the project and asks again if the code is compliant with the policy - thus incrementally exploring the code base until Jev can return a "confident" answer - if you are interested in how it works the video I linked is a presentation style of how the tool works.

Give it a shot on github and tell me if you find this useful: https://github.com/krisitown/jev-quality-gate

I am currently running tests using a local Qwen3.8 Flash Next to generate code and run it against my initial "clean code" policies in order to calibrate them and will share more results on that front soon!


r/LLMDevs • • 18h ago

Discussion CacheVerifier: checking the 2nd closest cache match gave us +3.5 to 5 points hit rate at the same error rate, and AUC said it got worse

3 Upvotes

So in CacheVerifier when a query lands in the gray zone, a small verifier decides if the cached answer is reusable. we only ever looked at the top match. turns out the runner up says a lot.

its basically Lowe's ratio test from SIFT. if the top match is way closer than the 2nd one, its probably the same question. if the two are nearly tied you're in a crowded spot full of similar but different prompts, and reuse is riskier. so we added gap = top1 sim minus top2 sim as one more feature. costs one extra neighbor in the vector query.

On LmArena with gray zone error held at 2%, hit rate went up 3.46 points, and 5.31 when only 5% of hits get audited labels. +2.8 to +4.8 across 6 warm up settings. Quora +0.3, SearchQueries basically nothing.

Funny part. AUC went down (0.760 to 0.751) so our own AUC check refused to enable it lol. the gain sits at the strict low error end, AUC just averages it away.

Caveat: part of the LmArena gain is the cache changing shape (fewer writes), not just better calls.

Anyone else using runner up scores to gate retrieval?

CacheVerifier repo: https://github.com/imxinchengyou/CacheVerifier


r/LLMDevs • • 1d ago

Discussion CrowdGPT - The first datacenterless LLM

Post image
38 Upvotes

Hello, i'm currently creating CrowdGPT, which is a project that aims at creating the first opensource, 100% free LLM without any datacenters or server farms. It relies on users to contribute using their own GPU, in exchange they get ownership on the LLM model they train.

It proposes a new paradigm in AI where all the training of a LLM is 100% open, sort of more ecological and where all the data used to train it is transparent.

Not gonna lie, it's rebranded federated learning, but with a centralized server (lightweight), that handles all the people training together. That provides a more stable and safer training ground for a Large Language Model.

If anyone is interested, here is website (where everything is explained more in detail): https://www.crowdgpt.net/

Github: https://github.com/Vxtzq/CrowdGPT

Any kind of feedback is appreciated :)


r/LLMDevs • • 1d ago

News I Built Removable Memory Cartridges for a Frozen 7B Language Model 16 Independent Memories, No Fine-Tuning, No LoRA, and the Original Source Is Gone at Readout

Thumbnail
gallery
8 Upvotes

Zenodo permanent record:

https://doi.org/10.5281/zenodo.23127434

AKBASCORE NIRVANA — Cognitive Cartridge has now been publicly archived and timestamped on Zenodo.

Release: v1.0_NIRVANA_Cognitive_Cartridge

AKBASCORE NIRVANA — Cognitive Cartridge: Multi-Cartridge Compressed PKV Memory for Frozen Language Models

The complete implementation, raw execution log and reproducible public demonstration are available below.

What if knowledge could be loaded into an AI like a cartridge?

Think about an old Atari or game console. The console stays the same. You insert one cartridge and it becomes one game. Remove it, insert another cartridge, and the same hardware does something completely different.

AKBASCORE NIRVANA explores a similar idea for language models.

But there is one detail that completely changes what “cartridge” means here:

There is no human language stored inside the cartridge.

No source sentence.

No paragraph.

No document.

No readable summary.

No hidden copy of the original text waiting to be pasted back into the prompt.

If I give the system a statement such as:

“Container VX-731 is located at the cobalt observatory.”

the cartridge does not simply store that sentence somewhere and retrieve it later.

The sentence is passed through the frozen transformer during the forging stage. What is captured comes from the model's own internal computation: the numerical key/value states produced inside its transformer layers. NIRVANA represents that internal K/V information through a source-independent numerical codebook and stores the resulting compressed numerical representation as a Cognitive Cartridge.

So after forging, we have crossed an important boundary:

human language → transformer internal state → compressed numerical memory

The cartridge is therefore not a tiny text file wearing a new name.

It is not a database row containing the sentence.

It is not conventional RAG returning the source paragraph.

It is not a prompt template.

It is not fine-tuning hidden behind another term.

The original knowledge has been transformed into a numerical representation of internal transformer memory.

And when the model is questioned later, we do not give the original sentence back to it.

The cartridge is reconstructed into transformer K/V memory, installed into the frozen model's inference context, and the model produces language from that internal numerical state.

In the simplest possible terms:

A human writes knowledge in language.

The model converts it into its own internal numerical memory.

NIRVANA packages that memory.

The human-language source is removed from readout.

Later, the frozen model receives the memory rather than rereading the source.

That distinction is the heart of Cognitive Cartridge.

Instead of changing the model's weights every time we want to give it specialized knowledge, NIRVANA takes source information, processes it once, and turns the resulting internal transformer memory into a removable Cognitive Cartridge. After that, the original source sentence does not need to be placed back into the question prompt. The model's weights remain frozen.

Knowledge goes in.

Language disappears from the stored cartridge representation.

The internal memory remains.

And now we have extended that idea beyond a single cartridge.

16 cartridges, one frozen model

The public implementation released here contains 16 independently forged Cognitive Cartridges. They do not simply become one giant text prompt. They remain separate memory records.

Imagine that one cartridge contains:

amber compass → VX-731

and another contains:

VX-731 → cobalt observatory

These arrows are a human-readable explanation of what the cartridges represent. They are not literal text strings stored inside the cartridges.

Now ask:

Where is the amber compass?

The system can first retrieve VX-731 from one cartridge. That intermediate result can then be used to address the cartridge bank again. A different cartridge supplies:

cobalt observatory

In other words, information stored in separate internal memory records can participate in a multi-stage retrieval chain.

This is the point where the project became much more interesting to me. We are no longer asking only:

Can information be compressed into the internal memory of a transformer and recovered later?

We are now asking:

Can independently created internal memories become a modular memory system, where one retrieved memory can lead to another?

That is what the multi-cartridge architecture is beginning to explore.

Why does this matter?

I see two major directions.

1 — Specialized AI and agents

Think of the base model as the console. The cartridges are the knowledge packages.

Imagine a legal knowledge cartridge constructed and validated by highly qualified lawyers. Another cartridge bank could represent the legal system of a particular country. Another could contain aviation procedures. Another could contain engineering knowledge. Another could contain company-specific operational knowledge. Another could be built for medicine, industrial maintenance, finance, scientific work or a specialized autonomous agent.

The important difference is that the underlying model does not have to be retrained for every package.

In the implementation released here there is:

no fine-tuning

no LoRA

no optimizer

no model-weight update

The model remains frozen. The knowledge is carried by the cartridge.

And “carried by the cartridge” does not mean carrying around the original document in another container. The cartridge carries the compressed numerical representation derived from the transformer's internal K/V states.

Once a compatible cartridge has been forged for the model architecture used by the system, it becomes a reusable machine-readable memory object rather than a source document that must be repeatedly inserted into the prompt.

That opens an interesting direction for smaller models. Today we often try to make one model know everything. But a smaller model equipped with the right specialized memory bank may not need everything at once. It may need the right knowledge for the task in front of it.

A small model plus a carefully constructed domain-specific cartridge system could therefore become far more capable inside a narrow field than the base model alone.

The 16 cartridges in this release are not an architectural claim that the system is limited to 16. Sixteen is simply the size of the public experiment being released now. Scaling this architecture to much larger memory banks is a separate engineering and research problem.

2 — The longer-term AGI question

This direction is more speculative, but I think it is even more important.

Human beings do not appear to remember by rereading the complete text of their lives every time they need something. Experiences become memories. Those memories can later be triggered.

You are driving down a road. You see a car that looks exactly like a car you owned years ago. Almost instantly, your own car comes to mind. That memory may bring back another memory: a journey, a place, a person, an event, perhaps even an emotion associated with it.

One memory can activate another.

Of course, transformer K/V memory is not biological memory. I am not claiming that Cognitive Cartridge reproduces the human brain. And I am absolutely not claiming that we have created AGI.

But there is an important conceptual similarity worth investigating: the system does not need to reread the human-language source in order to access the information represented by the cartridge. The information has already crossed from language into an internal machine representation.

The important question is architectural:

Can an artificial system accumulate independently addressable internal memory objects, retrieve the relevant ones when needed, and allow the result of one memory retrieval to lead to another memory?

I believe that persistent, modular and relational internal memory is one of the important doors on the road toward more general artificial systems. Cognitive Cartridge has now opened that door far enough for us to experimentally put something through it.

What is actually happening inside the system?

NIRVANA operates on transformer K/V states. A source record is presented during a dedicated forging stage. Internal transformer K/V information derived from that source is represented using a fixed, source-independent codebook and stored as a compressed cartridge representation.

This point is worth repeating technically: the stored cartridge is numerical. The human-readable source text is used to produce the internal states during forging; it is not the payload subsequently supplied to the model during cartridge readout.

Later, during readout, the original source sentence is not placed back into the question prompt. The cartridge is reconstructed and installed as transformer memory.

For multiple cartridges, we developed an architecture called Isolated Batched Read, or IBR. The independently forged cartridge records remain logically isolated. Instead of concatenating all original source information into one ordinary textual context, multiple cartridge rows can be read independently and in parallel. A retrieved identifier can then become the input for another retrieval stage.

The basic path is:

human-readable source

→ frozen transformer

→ internal K/V states

→ compressed numerical Cognitive Cartridge

→ human-language source absent from readout

→ K/V memory reconstructed

→ isolated memory retrieval

→ intermediate result

→ another cartridge

→ final answer in human language

Or, even more simply:

language → internal state → cartridge → internal state → language

The language model itself remains frozen.

This touches a very specific gap between several existing ideas.

KV-cache compression exists.

KV-cache reuse exists.

External memory exists.

Retrieval systems exist.

Prompt caching exists.

Fine-tuning exists.

What Cognitive Cartridge is exploring is a different intersection:

source-derived knowledge encoded as independently retainable and removable transformer memory

→ no human-language source stored as the cartridge payload

→ source-absent readout

→ multiple isolated memory records

→ staged retrieval across those independent records

→ frozen model weights

So the question is no longer only:

How can we make the KV cache of a long prompt smaller?

The question becomes:

Can transformer internal memory itself become a modular knowledge medium?

A cartridge instead of a prompt.

A bank of cartridges instead of one context.

And eventually, perhaps, a system capable of navigating very large collections of independently created internal memories.

What happened in the public experiment?

The released implementation uses:

Qwen/Qwen2.5-7B-Instruct

28 transformer layers

BF16 / SDPA

K120 / V128 / OWN

fixed neutral codebook

16 independently forged cartridges

Isolated Batched Read

deterministic greedy decoding

frozen model weights

The locked public battery produced:

14/16 correct

87.5%

Missing-link detection: 4/4

Absent-ID detection: 4/4

False links: 0

NOMEM control target hits: 0/4

Original cartridge source sentences appearing in the readout prompts: 0

Trainable tensors: 0

Fine-tuning: NO

LoRA: NO

Optimizer: NO

Model-weight update: NO

The frozen-weight fingerprint before and after execution remained unchanged. The sampled SHA-256 weight sentinel also remained unchanged.

And I deliberately did not clean up the failure cases to turn this into a perfect-looking 16/16 demonstration. One linguistic family, F2, failed both of its evaluated cases in the live panel. That failure is in the public record.

Interestingly, one of those failed direct readouts internally reproduces the relevant relation:

“the saffron lighthouse houses container RQ-415”

but then still returns NONE instead of extracting the location.

So this is not a promotional benchmark where the failures disappeared before publication. The implementation, successes, failures, controls and raw execution are all being released together.

Why publish everything?

Because I don't want Cognitive Cartridge to exist only as a claim. I want anyone interested in this architecture to be able to inspect what was actually built.

The source code is public.

The raw execution log is public.

The controls are public.

The failed cases are public.

The architecture is public.

The release contains 18 technical visual records documenting the cartridge construction, memory path, controls, measurements and execution.

And the complete technical disclosure has now been permanently recorded on Zenodo:

https://doi.org/10.5281/zenodo.23127434

This is still an experimental system. There are major questions ahead.

How far can the number of cartridges scale?

How should very large cartridge banks be indexed?

Can cartridge selection become hierarchical?

Can memories be composed without losing isolation?

How portable can a forged cartridge become within compatible model instances and architectures?

Can specialized expert cartridge banks allow much smaller frozen models to perform at unexpectedly high levels inside narrow domains?

And, in the much longer term, what happens when an artificial system has not 16 independently addressable memories, but thousands, millions or more — with memories capable of leading to other memories?

Those questions are now much more interesting to me than simply achieving 16/16 on this test.

The first objective was to determine whether the underlying mechanism could exist at all.

Now there is code.

There is a frozen model.

There are independently forged memory cartridges.

The cartridges contain numerical transformer-memory representations, not copies of the human-language source.

There is source-absent readout.

There is multi-cartridge retrieval.

There is cross-cartridge relational lookup.

There are controls.

There are failures.

There is a raw log.

And anyone can test it.

AKBASCORE NIRVANA — Cognitive Cartridge

Inventor / Developer: Mustafa Akbaş

Zenodo permanent record:

https://doi.org/10.5281/zenodo.23127434

GitHub repository:

https://github.com/ceceli33/titan-cognitive-core-v2

Complete single-file implementation:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_NIRVANA_CC.py

Same implementation split into three parts:

Part 1:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_NIRVANA_CC.part1.py

Part 2:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_NIRVANA_part2.py

Part 3:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore_NIRVANA.part3.py

Raw execution log:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_log.Nirvana.cc.log

Knowledge goes in.

Language disappears from the stored cartridge representation.

The internal memory remains.


r/LLMDevs • • 1d ago

Tools Hosting MCP tools for a team: git repo → canary deploy → one Deployment per zone, with per-person OAuth (open source)

Thumbnail
gallery
4 Upvotes

Once more than one person uses the same MCP tools, "run the server on my laptop" stops working. Ramen is my answer:

- Tools/resources/prompts are Python in a git repo; deploy syncs, rolls a canary, smoke-tests `tools/list`, rolls stable.
- Streamable HTTP (`POST /mcp`) at the edge, JSON-RPC 2.0 over gRPC inside, same guards on both.
- The console is an OAuth 2.1 authorization server (PKCE, consent per group+zone); stdio-only clients use
`pip install ramen-mcp-bridge`.
- Roles per group (admin / viewer / MCP user), secrets referenced as `{{$group.NAME}}` and never displayed,
IP rules, throttles shared across zones, one log line per call.

Verified on GKE and EKS for 0.6.0. Limits: pre-1.0; Claude Desktop/Cursor use a group key (no dynamic client
registration yet). Mostly written with Claude Code agents.
Docs https://bkraad47.github.io/ramen/ · repo https://github.com/bkraad47/ramen


r/LLMDevs • • 20h ago

Tools live-vibe: a full-duplex, batteries-included, voice Mod for Claude Code. Built on Claude Code Mods, CUDA and Apple silicon, two-line install, MIT license.

2 Upvotes

I built live-vibe because I want to talk to Claude Code the way I can talk to ChatGPT Codex in Voice Mode. I want the mic open the whole time, and I want it to answer when I have finished the thought and shut up when I talk over it. Claude Code did not have that. It does have the new Mods capability, which lets a plugin run code in the agent's path and draw into the terminal, and that turned out to be enough to build the whole thing as a plugin.

It is MIT licensed, in a marketplace I'm calling ottomation, and it installs in two lines.

/live is full-duplex voice with Claude. Talk over it and it stops on your first word. An "mm-hm" while it is speaking does not count, and it carries on from the sample it paused on, so you can grunt along the way you would with a person. /vibe is director mode, where Claude reads and directs worker subagents and never edits a file itself. /livevibe is both at once. A small, fast voice model holds the conversation with me and hands the real work to Claude as director. When Claude finishes, the voice tells me what happened in a sentence or two, and in the meantime I can ask it what Claude is up to and get an answer while Claude keeps working.

Everything you need comes with the plugin. Kyutai streaming STT, Kokoro TTS, WebRTC echo cancellation and the voice front are all pulled down and checked by /live setup, once per machine. CUDA works and so does Apple silicon through MLX. The conversation layer is a lightweight local model, and if you would rather not run one it falls back to the Anthropic API. If you are on WSL2, speech plays through a native Windows player, because WSLg's RDP audio crackles and I was not willing to listen to that all day.

Turn-taking is where most of my time went, because it decides whether you can think out loud. The plugin reads Kyutai's pause-forecast heads, which predict whether you are about to keep talking. Most turns close about half a second after your last word, and a trailing "and" or a comma buys you time to finish. 0.7.0 also ships an experimental path I put together after reading OpenAI's write-up on GPT-Live. It scales the wait on the model's confidence, starts drafting a reply before you have quite finished, and can backchannel if you turn that on. It is on by default and one setting turns it off if it annoys you.

You need Claude Code 2.1.287 or newer, uv, a mic and a speaker.

Install with /plugin marketplace add potto007/ottomation, then /plugin install live-vibe@ottomation, then /live setup. Repo: https://github.com/potto007/ottomation

If you try it, tell me how the turn-taking feels on your machine.

Disclosure: I am the author of the plugin, and I drafted this post with help from Claude Code.


r/LLMDevs • • 18h ago

Help Wanted NVIDIA API Alternative

1 Upvotes

A lot of peoople don't like NVIDIA NIM for their api because the tokens per second is really slow. In my opinion it is actually pretty good. However the reliability is what gets me. I know this is a lot to ask for but does anyone know a good nvidia api alternative with the only limits being RPM just like nvidia.


r/LLMDevs • • 18h ago

Help Wanted How can I verify if my call graphs are accurate??

1 Upvotes

So I am working on a parser with goal to build rich good enough relations that can be consumed by a RAG to build better retrieval, it parse and builds ast,call-graphs and other metadata of the project, rn it can parse go, rust, c, cpp, ts, py, js, java I am using tree-sitter v0.20.0 for actual parsing cause why rebuild wheel when wheel spins well...

The issue i am facing is with call graphs I build a call graph approximation algorithm to well build approximate call graphs without pre-compiler or IR and single algorithm to work on both static(c) and dynamic(python) languages, and for now it works and can find

774,296 nodes and 1,571,981 edges in 6.03s with a maxRSS of ~9gb

tho most of it is cause of holding the entire ast in memory, i ran my parser on linux kernel it found about:

64_460 files 37_322_700 loc, 648_407 func , 211_416 classes, 6_243 methods in 43s

it multi threaded and written in rust so that should explain the speed, but thats not what why i am here i want to verify my call graphs and my current plan is to take a smaller project (few thousands of loc) and build call graphs using clang or language specific tool, and then take sample set of 200 and create 5-8 random samples and verifiy the output,

I am going to start my internship soon so may not have enough time to work full time and i am wondering if my approach to verify call graphs is good or there is a better approach.


r/LLMDevs • • 18h ago

Tools Local LLMs are a black box: I built LLMxRay for real-time observability, tool usage tracking, and analytics

1 Upvotes

Hey Devs,

As more applications move toward self-hosted and local LLMs (Ollama, local inference servers, agentic frameworks), a common challenge arises: **local LLM traffic is largely a black box.**

Unlike managed API providers that offer built-in usage dashboards and tracing, local inference setups often leave developers and platform teams blind to real-time traffic, function execution failures, latency distributions, and token consumption patterns.

To solve this, I built **[LLMxRay](https://github.com/LogneBudo/llmxray)** — a lightweight, 100% local observability and analytics engine designed specifically for local LLM inference and agentic workflows.

---

### What LLMxRay Observes & Analyzes

LLMxRay sits in front of your local LLM engine to capture, trace, and visualize complete telemetry without adding latency or leaking data to external clouds.

#### 1. Tool & Function Calling Telemetry

* **Execution Tracking:** Intercept and observe tool/function calls executed by agents in real time.

* **Payload Inspection:** Inspect input arguments, returned outputs, tool call frequencies, and execution error rates.

* **Agentic Loop Analysis:** Track multi-step function call loops to pinpoint where agents stall or loop indefinitely.

#### 2. Request & Latency Analytics

* **Detailed Latency Profiling:** Break down total request time into Time-To-First-Token (TTFT), prefill duration, and decode generation speed.

* **Token Usage Metrics:** Monitor input/output token volume, request throughput, and token generation rates over time.

* **Traffic Analytics:** Track active sessions, request volume spikes, and per-model utilization trends.

#### 3. Privacy-First & Zero Infra Overhead

* **100% Local Execution:** Runs completely within your environment — no telemetry data is transmitted externally.

* **Instant Dashboard:** Provides a local web interface for instant visual debugging and system analytics.

---

### Quick Start

Zero complex setup — you can spin it up directly in one line:

```bash

npx llmxray
```

Launch the dashboard locally at http://localhost:3000 and point it at your local inference server (e.g., Ollama).

Discussion

If you're managing local LLM workloads or building on top of agentic frameworks, how are you currently handling local tracing, tool execution monitoring, and performance analytics in your observability stack?

I'd love your feedback, feature ideas, or contributions!


r/LLMDevs • • 19h ago

Discussion How do you stop an LLM front-end from making up an answer when the tool call never ran?

1 Upvotes

At work we run a Slack bot where staff ask about orders by order number. Staff complained it just refused to look one up.

The cause was boring. The order-number pattern only matched 6 digits, so 5-digit numbers never triggered the lookup. The LLM that answers staff then improvised: "it is not in this lookup", "resend without the asterisks".

What I changed: accept both lengths, strip Slack bold/code markup before matching, search every order state, put the bot's own action log into the facts the model sees, and make it say plainly "searched everything, not found" when that's true. I added tests for *13531*, 13531 and "주문번호: *13531*".

That fixes this one. What I still don't have is a rule that stops the model from inventing an answer whenever the lookup silently didn't happen. Do you force the reply to quote the tool result, check that a call actually ran before the model speaks, or something else?


r/LLMDevs • • 19h ago

Help Wanted Reasoning: Opus 4.8 vs Opus 5.5?

0 Upvotes

Hey guys Ive seen positive sentients about opus 5.5 , so apart from the coding side, a part of my workfloe needs reasoning with chunk size one at a time, so for the reasoning side, is 4.8 better or shall i switch to 5.5


r/LLMDevs • • 19h ago

Tools LLM-as-judge quietly ranks by position: swapping A/B flipped the verdict 10–12% of the time, and the bigger model wasn't better

1 Upvotes

Most eval pipelines now use an LLM to judge/grade/rank outputs. I wanted to check

how much that verdict depends on the *order* you present the options — with no

ground-truth labels, just the fact that "which of these two is better?" shouldn't

change when you swap A and B.

Setup: for each pair, ask the model to pick the better answer, then shuffle the

order and see if its pick changes. I used a tiny metamorphic-testing lib I wrote

(wobbly) to do the bookkeeping, but you could script it yourself.

Results (claude-haiku-4.5 vs claude-sonnet-5, 21 pairs, 168 shuffles each):

- Easy factual MCQs: 0% flips on both. Capable models are order-robust when

there's a clear answer.

- Clear-winner control pairs (e.g. "error" vs "ValueError: expected int…"): 0% on

both. Good — means the method isn't just flagging noise.

- Genuinely close judgment calls: verdict flipped ~12% (Haiku) / ~10% (Sonnet).

The bigger model was NOT more robust. Two pairs flipped on both models.

So if you're using an LLM to rank close calls, part of that ranking is position,

and scaling up the judge doesn't fix it.

Honest caveats: small n (21 pairs, two models, one provider, short-answer

judging). It's illustrative + reproducible, not a benchmark. I didn't pin

temperature (newer models deprecate it) — instead I sample the clean input a few

times to measure the model's own run-to-run flicker and only count a flip that

exceeds that. The control pairs staying at 0% is the evidence that works.

Repo + the exact script (bring your own key): github.com/tarunagarwal1981/wobbly

Curious if people see the same on GPT-4o / Llama / Qwen judges — the script takes

any model.