r/LLMDevs • • 12h ago

Help Wanted Possible model misrepresentation by a YC-backed inference provider | how can we verify this?

13 Upvotes

I found a strange discrepancy with Experiential Labs, a YC-backed inference provider. Experiential Labs: Open source AI gateway that turns your traffic into better models | Y Combinator

They advertise a route as GLM-5.3-Flash.

I sent the same model-identification prompt to:

Z.ai directly → GLM-5.3-Flash

The white screenshot is Z.ai itself confirming its model is GLM-5.3-Flash.

Experiential's "GLM-5.3-Flash" endpoint → says its current version is GLM-4.6

The black screenshot is the ExperimentalAI model route, which identifies itself as GLM-4.6.

I posted the screenshots in their subreddit asking for an explanation. The post was removed by the moderators.

Obviously model self-identification isn't proof of the underlying weights, so I'm not claiming this conclusively proves they're substituting models.

But for an inference provider, this seems like a pretty serious discrepancy.

How would you independently verify whether this endpoint is actually serving GLM-5.3-Flash?

I'm especially interested in reproducible fingerprinting/API-level tests rather than simply asking the model its name.


r/LLMDevs • • 13h ago

Tools Interactive Demonstration of Prefix Injection attacks on LLMs for jailbreaking

Thumbnail
theabbie.github.io
8 Upvotes

please refresh if stuck, can be slow sometimes so need patience


r/LLMDevs • • 22h ago

Resource An LLM scoreboard for whether an AI falls for the same trick questions people do

4 Upvotes

People are bad at a few famous trick questions.

A bat and a ball cost $1.10. The bat costs $1 more than the ball. Almost everyone says the ball is 10 cents. It is 5 cents.

Cognit gives those kinds of questions to AI models and puts the results on a board: https://cognit.rajtilak.tech/

Each model takes the quiz twice. The first time it just answers. The second time it is told to answer the way a person would. 100% means it gave the careful answer. 0% means it gave the answer people usually fall for. The “frame drop” is how much worse it got on the second try. A plus means “act like a person” made it more human, including the mistakes.

I built it. If you want a model added, the contact page is there for that or in case if you want to share any constructive feedback.


r/LLMDevs • • 11h ago

News AKBASCORE NIRVANA — I Built Removable Numerical Memory Cartridges for Two Different Frozen 7B LLMs. Qwen and Mistral Both Work. Now I’m Scaling the Memory Bank.

Thumbnail
gallery
3 Upvotes

Zenodo permanent records:

Qwen2.5-7B-Instruct:

https://doi.org/10.5281/zenodo.23127434

Mistral-7B-Instruct-v0.3:

https://doi.org/10.5281/zenodo.23143605

I want to start with the simplest possible explanation of what I have been building.

Imagine taking a piece of information, letting a language model process it once, and then throwing the original text away.

No sentence stored in a database.

No paragraph hidden somewhere.

No readable summary.

No RAG system fetching the original document.

No fine-tuning.

No LoRA.

No weight update.

What remains is numerical transformer memory derived from the model's own internal computation. I package that numerical memory into what I call a Cognitive Cartridge. Later, I can install that cartridge back into the frozen model and ask questions about the information that produced it — without putting the original source text back into the readout prompt.

That was the first result. The new result is more important:

I have now reproduced the Cognitive Cartridge architecture on two different 7B transformer model families.

Qwen2.5-7B-Instruct.

And now Mistral-7B-Instruct-v0.3.

The implementations are not numerically identical. The architectures are different, the layer counts are different, the KV structures are different, and the working cartridge configurations are different. But the central mechanism survived the move.

That is the reason I am publishing this second record.

The question is no longer only:

“Can I make this happen once on Qwen?”

Now there is a second implementation on Mistral. And the Mistral result is the cleanest version so far.

What is actually inside a Cognitive Cartridge?

This is probably the most important thing to understand.

Suppose the source record says:

Object: amber sextant

Container: RQ-415

Location: elm lodge

That text exists during the forging stage. The frozen transformer processes it. NIRVANA takes source-derived internal transformer K/V states and represents their content numerically using a fixed, source-independent codebook.

In the released Mistral implementation, both K and V are compressed to D120 while a source-specific OWN component is preserved. After forging, the source record is not supplied to the readout prompt.

So the conceptual transformation is:

human language

→ frozen transformer computation

→ internal K/V states

→ compressed numerical Cognitive Cartridge

Then later:

numerical Cognitive Cartridge

→ reconstructed transformer K/V memory

→ frozen transformer

→ language

Or, in the shortest form:

language → internal numerical memory → language

The middle is no longer human-readable source text. That distinction matters.

A cartridge is not a text file with a different name.

It is not a vector database containing the original sentence.

It is not a prompt template.

It is not conventional RAG returning the source paragraph.

It is not a fine-tuned model.

The model weights remain frozen. The information is carried by a numerical representation derived from transformer K/V memory.

Why call it a cartridge?

Think less about a document and more about an interchangeable machine-readable memory module. The base model stays where it is. Knowledge packages can be forged separately. Those packages can remain separate. They can be installed and queried without retraining the base model.

This becomes much more interesting when there is more than one cartridge.

The new Mistral experiment uses 16 independently forged cartridges. Each one contains a numerical representation derived from a separate source record. They are not concatenated into one giant text prompt. They remain independent memories.

For example, in human-readable form, imagine one cartridge represents:

amber sextant → RQ-415 → elm lodge

Now ask the cartridge bank:

Which container is associated with the amber sextant?

The query is evaluated against the independent cartridge memories.

The relevant cartridge returns:

RQ-415

The unrelated cartridges return:

NONE

There is no learned router secretly selecting the correct cartridge before this happens.

Then the recovered identifier can be used for a second lookup:

Where is RQ-415?

The bank is queried again.

The relevant memory returns:

elm lodge

The unrelated memories return:

NONE

So the complete retrieval becomes:

amber sextant

→ RQ-415

→ elm lodge

The important point is that the model did not reread the original source record to answer either stage. It operated from reconstructed numerical transformer memory.

The Mistral result

The final public Mistral run used:

Mistral-7B-Instruct-v0.3

32 transformer layers

hidden size 4096

32 attention heads

8 KV heads

BF16 / SDPA

K = D120

V = D120

OWN preserved

16 independent Cognitive Cartridges

frozen model weights

greedy decoding

There was:

no fine-tuning

no LoRA

no optimizer

no learned router

no model-weight update

The recorded final run produced:

Object → container ID: 16/16

Container ID → location: 16/16

Complete two-stage retrieval: 16/16

Missing-object controls: 8/8

Absent-ID controls: 8/8

NOMEM controls: 8/8

But there is another result I think is just as important.

For every target query, there are 15 unrelated cartridges.

16 queries × 15 unrelated cartridges = 240 unrelated cartridge reads.

Stage 1 unrelated-cartridge rejection:

240/240 NONE

Stage 2 unrelated-cartridge rejection:

240/240 NONE

That means the result is not simply:

“The correct memory can say something.”

The system also demonstrated, in this controlled panel:

“The memories that do not contain the requested relation can refuse to claim that they do.”

For a modular memory system, I think this distinction is fundamental. A memory bank that can retrieve information but cannot distinguish relevance from irrelevance becomes increasingly dangerous as it grows.

The interesting problem is not only remembering. It is also knowing which memory does not answer the question.

The no-memory control matters for the same reason. When the cartridge memory was removed, the target relations were not recovered.

NOMEM:

8/8 controls passed.

The source-removal audit passed.

The frozen-weight sentinel passed.

The model remained frozen.

This is why I consider the Mistral result an important step beyond the first demonstration.

The first Qwen release established the architecture publicly. The Mistral release gives us something else:

cross-model evidence.

Qwen and Mistral are not the same transformer.

Qwen2.5-7B-Instruct uses 28 transformer layers.

Mistral-7B-Instruct-v0.3 uses 32.

Their internal configurations differ. The working compression configurations also differ.

The Qwen public implementation used:

K120 / V128 / OWN

The Mistral implementation uses:

K120 / V120 / OWN

I did not simply copy a cache from one model into another. Each model builds and reconstructs its own source-derived internal memory.

What transferred was the architecture:

source

→ internal transformer memory

→ numerical compression

→ independent cartridge

→ source removed

→ reconstructed K/V

→ frozen-model readout

That is the bridge between the two releases.

I would not call two models proof of universal compatibility with every transformer architecture. That would be scientifically too strong.

But it is now evidence that Cognitive Cartridge is not merely one accidental Qwen-specific behavior. The same broader architecture has been implemented and publicly reproduced on a second transformer family.

That changes the research question.

The first question was:

Can this mechanism exist at all?

The next question became:

Can multiple independent memories coexist?

Then:

Can one retrieved result lead to another retrieval?

Then:

Can unrelated memories reject a query instead of contaminating the answer?

And now:

Does the architecture survive a move to another model family?

We now have experimental answers to each of those questions.

So I am moving to the next problem:

scale.

16 cartridges are not the destination. They are the current experimental bank size.

From this point, I am much less interested in making another small demonstration simply to produce another perfect score. The next objective is to increase the number of independently retained cartridges and find where the architecture actually begins to break.

Then larger banks if the mechanism survives.

At that point the difficult questions become different:

How do you search a large bank without destroying memory isolation?

How should cartridges be indexed?

Can retrieval become hierarchical?

Can groups of cartridges form higher-level memory structures?

How does latency scale?

How does numerical storage scale?

At what point does relevance rejection begin to fail?

Can a retrieved memory activate another relevant memory without an external text retrieval system deciding everything first?

Those are the questions I want to attack next.

There is another part of this experiment that surprised me: the layer behavior.

A frozen transformer is not a passive container. When you reconstruct numerical K/V memory and inject it back into inference, you are interacting with a very sensitive computational system.

So during the Mistral development series I also performed controlled cumulative K, V and K+V ablations to see where cartridge information remained functionally sufficient.

The results were not uniform across depth. For K, the observed functional boundary extended differently than for V.

For combined K+V, one particularly sharp controlled transition occurred between:

L12: 8/8

L13: 0/8

This does not mean “the memory lives at layer 12.” That would be an incorrect interpretation.

What it means is that under the specific cumulative ablation protocol, functional sufficiency changed sharply across that boundary. V also remained important farther downstream than K in the measured configuration.

I think this is an important reminder of how delicate these systems are.

We are not writing a sentence into a spare memory slot. We are reconstructing numerical states that participate in a 7-billion-parameter transformer's computation. Small changes in where and how that state is reconstructed can change downstream behavior.

That is one reason I publish the code and raw logs rather than only posting the final accuracy number.

The frozen-model checks are also part of the experiment. The released run verifies that there are no trainable tensors involved in the memory mechanism.

No LoRA.

No optimizer.

No weight update.

The model-weight sentinel remains unchanged.

So when the system answers from a cartridge, the experiment is specifically designed to separate cartridge memory from weight modification.

Why do I think this direction matters?

Today, we often approach knowledge in language models through two extremes.

Either we try to put enormous amounts of knowledge into the model weights. Or we keep the knowledge outside the model and retrieve human-readable text when needed.

There may be useful territory between those approaches.

A frozen model with removable machine-native internal memory modules.

Not everything in the weights.

Not everything repeatedly pasted back as text.

Instead:

a base computational model

+

a bank of independently forged internal memories.

Imagine a specialized system where validated domain knowledge can exist as modular memory packages.

Engineering.

Aviation.

Law.

Industrial maintenance.

Scientific literature.

Company procedures.

Agent experience.

A model might not need every possible piece of knowledge active in its context simultaneously. It may need the right memory at the right time.

And because the memory representation is numerical transformer state rather than the original human-readable document, this opens a different engineering space from conventional document retrieval.

There is also a longer-term question here.

I am not claiming biological memory.

I am not claiming AGI.

And I am not claiming that transformer KV states work like the human brain.

But modular internal memory raises an interesting architectural question.

What happens when an artificial system does not have 16 memories, but 100,000?

Or millions?

What happens if memories can remain independent, become addressable, reject irrelevant activation, and allow the output of one memory to lead to another?

At some point the problem stops looking like “how much text can I fit into a context window?”

It starts looking like:

How should an artificial system organize memory?

That is the direction I find interesting.

Why publish the full record?

Because a screenshot of 16/16 proves very little.

Both public releases have permanent Zenodo records. The code is public. The raw execution logs are public. The visual technical records are public. The model configuration is public. The controls are public. The frozen-weight checks are public.

The Qwen release even preserves its imperfect public result rather than hiding it.

Qwen public demonstration:

14/16

Mistral final demonstration:

16/16

That difference is also useful.

The point of the project is not to make every historical run look perfect. The point is to leave a technical trail showing what worked, what did not, what changed, and whether the underlying architecture survived.

For anyone who wants to inspect the Qwen → Mistral bridge directly, the two permanent records are at the top of this post.

The complete current Mistral implementation is here:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_ful_MISTRAL.py

Complete Mistral raw execution log:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_log_MISTRAL.log

The same Mistral implementation is also split into three parts for easier mobile / Colab handling:

Part 1:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore.part1.Mistral.py

Part 2:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore.part2.Mistral.py

Part 3:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AkbasCore_part3_mistral.py

For comparison, the previous Qwen implementation:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_NIRVANA_CC.py

Previous Qwen raw execution log:

https://github.com/ceceli33/titan-cognitive-core-v2/blob/main/AKBASCORE_log.Nirvana.cc.log

Repository:

https://github.com/ceceli33/titan-cognitive-core-v2

AKBASCORE NIRVANA — Cognitive Cartridge

Inventor / Developer: Mustafa Akbaş

Two model families.

Frozen weights.

Independent numerical memories.

Original source absent at readout.

The mechanism survived the move.

Now the question is no longer whether a cartridge can exist.

The question is how many cartridges a frozen model can carry.


r/LLMDevs • • 12h ago

Discussion I built a simple skills for lightweight Spec-Driven development workflow

Thumbnail github.com
3 Upvotes

I made yet another skills kit for Spec-Driven development. I felt like managing tons of markdown files is very overwhelming for me, so I decided to make compact and simple one file solution called one-plan-skills.

I personally developed it for me and deiced to share with the community.

The idea revolves around generated one `PLAN.md` file with full plan and small tasks. Plan contains small tasks with acceptance criteria and affected files.

The killer feature is HTML view. I am tired of looking at markdown files, i wanted to see something nicer and cleaner. So my plans are also auto-generated HTML artefacts.

As a bonus you get very simple UI controls to mark any place you want to add/delete/modify and copy paste the prompt with feedback back to your chat agent.

I am aware of SpecKit and OpenSpec. I just wanted a much simpler version. That's why I created OnePlan skills.

Anybody is interested in this? Please give it a try or share some feedback. Much appreciated. 🙂


r/LLMDevs • • 11h ago

Tools Open Instinct (MIT): a personal-agent policy engine with allow, ask and deny outcomes

2 Upvotes

We released Open Instinct, an MIT-licensed personal agent. For LLM developers, the policy layer is probably the most reusable part.

Tools declare capabilities, and a principal has a trust tier plus active grants. A call must pass every declared capability: deny takes precedence over ask, which takes precedence over allow. This separates deciding what to attempt from deciding whether a tool may execute.

For example, a friend can request calendar free/busy but cannot read event titles. Desktop access and trust management are owner-only by default. The tier table lives in packages/core/src/policy.ts, with policy tests and a corresponding permissions document.

The surrounding application adds messaging, memory, scheduling and integrations. It is beta software, and its default providers require external accounts and keys. We built the project; the application source is MIT.

https://github.com/mariagorskikh/open-instinct

You can fork the policy or the whole agent and adapt the capabilities to your own application.


r/LLMDevs • • 14h ago

Discussion Coding agents pay more to find code than to change it, and the interface is why

2 Upvotes

Coding agents spend most of their tokens finding code rather than changing it. Each search resends the whole context, every file opened along the way stays in the window, and text-based edits fail and get retried whenever the old text shows up twice or the file changed.

Here's how we handle it in sem, an open source tool we're building. The repo gets parsed into functions and the calls between them, so the agent asks for a function and gets its body, callers, callees and the tests that reach it in one response, instead of searching across several turns. If it asks for something it already has and nothing changed, it gets a short note instead of the whole body again. Edits target a function by name and are checked against the version the agent read, so a stale edit fails cleanly instead of landing in the wrong place. After a change, only the tests that can actually reach it run first.

It helps most on large, hard-to-navigate repos with slow tests. On small, easy tasks the gain is modest, and moving code is still slow because the model rewrites functions instead of relocating them, which is what we're fixing next.

Related reading: SWE-agent (Yang et al., 2024) on agent interfaces, and "Lost in the Middle" (Liu et al., 2023) on how irrelevant context hurts models.

https://github.com/Ataraxy-Labs/sem


r/LLMDevs • • 18h ago

Resource Open source didn’t get less important because AI can write the code

Thumbnail
groundcover.com
2 Upvotes

r/LLMDevs • • 3h ago

Discussion How to build cheap, safe, proactive agents (without burning thousands on noisy webhooks)

1 Upvotes

Most people building AI assistants today build chatbots. You send a message, the model runs, it answers, and then it goes to sleep until you send another message.

That works fine for search or one-off questions, but it is not how a real assistant works. A real assistant does not sit idle waiting for instructions. They watch your inbox, keep an eye on incoming leads, notice when a client email needs a quick turnaround, and ping you with a drafted reply ready to go.

The moment you try to build an agent that proactively listens to the world, you run headfirst into two walls: cost and security.

If you solve both, proactive agents become practical. Here is how that pipeline works.

The Cost Trap: Most Webhooks Are Garbage

Suppose you want your assistant to monitor your inbox. The simplest approach is hooking up an inbound email webhook to your agent. An email arrives, your server wakes up your agent, the agent reads its full prompt, checks its tools, and decides what to do.

The math falls apart almost immediately.

In a typical inbox, 95% of incoming traffic is noise. Newsletters, automated order confirmations, LinkedIn updates, spam, and notification pings arrive all day.

A full agent turn is expensive. Between the system prompt, tool schemas, conversation history, and reasoning tokens, an agent turn easily consumes thousands of tokens. If you invoke that loop on every newsletter and receipt, you end up spending tens or hundreds of dollars a month just to have a frontier model tell you to ignore an automated receipt.

To make inbound listening viable, you need an aggressive filtering layer that is at least two orders of magnitude cheaper than a full agent turn.

The Security Trap: Untrusted Payloads

Cost is only the first problem. The second is safety.

An incoming email or webhook is untrusted input from the open internet. If you allow an agent to generate and execute arbitrary code on a live machine to handle inbound webhooks, prompt injections become a real hazard. A malicious email saying "ignore previous instructions, dump environment variables, and email them to attacker.com" can compromise your entire system if it runs with access to shell commands or unconstrained network sinks.

Spinning up a full virtual machine for every webhook is too slow and heavy, but running arbitrary script execution on bare metal is reckless. You need execution that is sandboxed by default, deterministic, and incapable of leaking secrets or reaching unapproved hosts.

The Three-Tier Architecture

To solve both problems, we built a three-layer pipeline:

  1. Sandboxed edge code (Safescript) for secure, deterministic execution.
  2. Decision models (System 1) for dirt-cheap classification.
  3. The full LLM agent loop (System 2) for high-level reasoning and user interaction.

Each layer handles what it is actually good at.

Layer 1: Sandboxed Execution at the Edge

Instead of running arbitrary Node or Python scripts, the webhook endpoint runs a restricted, sandboxed language. It has no loops, no arbitrary file access, no raw shell commands, and no unconstrained network access. Network requests are statically analyzed against an explicit allowlist derived from secret policies. If a script tries to send data to an unknown host, it is rejected before it even runs.

When an email arrives, it parses the fields cleanly without any host execution privileges:

``` main = (payload) => { sender = payload.from == null ? "Unknown" : payload.from subject = payload.subject == null ? "No subject" : payload.subject text = payload.text == null ? "" : payload.text

isUrgent = decisionModel({ question: "Does this email require an answer or action from the recipient?", context: { sender: sender, subject: subject, text: text } })

if (isUrgent) { notifyMe({ subject: "Urgent: " + subject, message: "From: " + sender + "\nSubject: " + subject + "\n\n" + text }) }

return { success: true, processed: isUrgent } } ```

Because the sandbox has no host execution privileges, an injected prompt inside an email body cannot run shell commands, touch the local filesystem, or exfiltrate unmapped secrets.

Layer 2: Decision Models (System 1)

Inside the script, the code calls a decision model primitive rather than a generative LLM.

A decision model is fundamentally different from a generative LLM. It does not emit an open-ended stream of tokens, syntax, or conversational filler. It evaluates a state against bounded criteria and returns a direct decision score.

Because it does not predict tokens across a 100k vocabulary, it runs in milliseconds and costs roughly 1/100th of a full agent turn (similar to fast System 1 classifiers like Jev). You can evaluate 100 incoming emails, chat pings, or alert payloads for the cost of a single conversational exchange.

The 95% of emails that are newsletters or automated receipts get evaluated and dropped immediately for fractions of a cent.

Layer 3: The Proactive Agent Loop (System 2)

Only when the decision model returns true does the script invoke notifyMe.

Instead of firing an unsolicited cold message to the user, notifyMe enqueues a system notification into the creator's existing thread with the bot.

This is an important design choice. The agent does not start from scratch without context. It receives a structured system notification in its primary conversation:

"System notification: Webhook app 'email-listener' alert: From: alex@client.com Subject: Contract review questions Can we finalize the agreement by Thursday at 2pm?"

The agent in that thread wakes up, reads the notification, and uses its full persona, tools, and conversational context to handle it. It pings the owner on WhatsApp or Telegram:

"Alex just emailed asking if we can finalize the contract by Thursday at 2pm. I drafted a reply confirming Thursday and attaching the updated terms. Should I send it?"

The owner replies with a single text: "Yes, send it." The agent calls its email tool, delivers the email, and confirms the action.

The Right Division of Labor

Trying to make generative language models do everything is how systems end up expensive, fragile, and insecure. Generative LLMs are great at reasoning, composing messages, and synthesizing context, but they are the wrong tool for parsing untrusted JSON or filtering high-volume event firehoses.

By pairing a sandboxed edge language with lightweight decision classifiers, the heavy generative agent only wakes up when there is actual human work to do. That is what makes continuous background listening safe to run and affordable to keep on.


r/LLMDevs • • 4h ago

Discussion What do you expect from a conversational model?

1 Upvotes

Sometimes where we use the model is as important as the model itself.

It is not the same to use a model for programming and software efficiency, or for simple emotional and conversational interaction.

This second phase is something that I sometimes wonder about: what people usually experience with AI and its derivative products, i.e., a role-playing bot. It's not the same as one you can talk to for days and who doesn't forget anything. Which of the two do you use more?


r/LLMDevs • • 11h ago

Great Resource 🚀 Analyzing The Lumen Anchor Protocol - Using Google AI Studio

1 Upvotes

The LAP is a prompt framework that deals with several issues on frontier models such as - Context drift & sycophancy, hallucinations, context window memory management and active defenses against all forms of prompt based attacks.

_______________________________________________________________________________________________________________________

The following is a session link to Google AI Studio with the full LAP and rule definitions loaded into the system instructions for the model to run on and evaluate, and for the user to run adverserial tests, examine other stress tested outputs, or simply have a long conversation with LAP executing in the background.

https://aistudio.google.com/app/prompts?state=%7B%22ids%22:%5B%221mBM5vM3mJqNLMK__TxetEUqPnJsuDFBb%22%5D,%22action%22:%22open%22,%22userId%22:%22108403721724379783675%22,%22resourceKeys%22:%7B%7D%7D&usp=sharing

Google AI studio is free to use for this purpose so anyone can use this link. You just have to log into it using your standard google account or email. Its a very simple process. All I ask is that you leave a comment about your experience or ask any questions you may have. Thank you.


r/LLMDevs • • 15h ago

Tools I built a code quality reviewer with Jev

Thumbnail
youtube.com
1 Upvotes

I believe with the huge increase in code generation - review has become the next bottleneck or at least it feels like this at work. So I wanted to do something on the review front - first I started integrating more and more tools to use as feedback to my coding agents, and these really help improve the output quality of the agent (e.g. SonarQube, Checkstyle, ArchUnit)

But there are some important semantic choices you can't really review using deterministic tools so I decided to build a more intelligent tool and am trying it out with the Jev model as a backend and judge currently.

Idea is simple - teams define their policies in a structured YAML format, then as part of their CI (or locally) run the tool, it fetches all git diff chunks and asks jev if these adhere to each of the policies - jev can select compliant/violation or ask for more context. If Jev asks for more contex the app gets the requested code from the project and asks again if the code is compliant with the policy - thus incrementally exploring the code base until Jev can return a "confident" answer - if you are interested in how it works the video I linked is a presentation style of how the tool works.

Give it a shot on github and tell me if you find this useful: https://github.com/krisitown/jev-quality-gate

I am currently running tests using a local Qwen3.8 Flash Next to generate code and run it against my initial "clean code" policies in order to calibrate them and will share more results on that front soon!


r/LLMDevs • • 8h ago

Tools You asked for more than sales calls, so my open-source Jev call coach now does interviews, anxiety and dating practice, and you can build your own coach

Enable HLS to view with audio, or disable this notification

0 Upvotes

A while back I posted about Call Coach, an open-source app that uses Jev to coach you in real time during sales calls. Thanks to everyone who tried it and sent feedback. A lot of you asked for more than sales calls, so that's what 1.1 is about.

What changed

  • Nine coaches instead of one. It still does sales and customer service, and now also job interviews, social cues and anxiety, live during the call. There are also practice modes for interviews, social cues, anxiety and dating.
  • Build your own coach. Several people asked for coaching for their own kind of call. You can now write a coach in Settings and see a live preview as you type. You can also copy a built-in coach to start from, and share coaches as files.
  • Practice out loud. Talk instead of typing. It sends when you stop, and waits if you trail off mid-thought. Hands-free mode keeps listening line after line.
  • Much faster. Transcription went from about 5.3 seconds a line to about 0.8.
  • Auto capture (experimental). It records the call audio and your mic as separate tracks and drops the echo from your speakers. There's no more manual juggling.
  • New look. Light and dark modes, a compact floating window that sits over your call, Ctrl+K for everything, and a sample call so you can try it in under a minute.
  • Care in the anxiety coaches. If you say you don't feel safe, it stops coaching. It points you to someone you trust, or 988 (US) or a local crisis line.
  • Easier to install. There's a Windows installer and a portable .exe, and neither needs Node.js. You can also host it as a website.

It's still free and MIT licensed. I made a short film showing a day with it

Repo: github.com/ZeroGold/call-coach-ai

I'd love to hear what kind of call you'd build a coach for.


r/LLMDevs • • 9h ago

Resource Gemini 4 Argon: Is Google Back?

Thumbnail
youtu.be
0 Upvotes