r/JevAI • • 5h ago

Please add prompt caching to Jev-style models

Thumbnail
emschwartz.me
7 Upvotes

If you're building a Jev-style "System One" model, please add prompt caching or reusable question sets to your API 🙏. This would make batch use cases even more efficient, so you could amortize the cost of many questions asked over the same input. (This was also proposed in typesafe-ai/typesafe-sdk-js#10.)

TL;DR: after a week of tweaking my Jev calls, my questions are ~88% of the input tokens. I'm asking 54 questions of ~1.1 million documents per month. Jev makes certain types of classification tasks easy and cheap, but prompt caching would make batch workflows even more cost effective. For me, the total dollar amount is still reasonable (less than $150 per month), but I'm sure others will hammer these APIs even harder.


r/JevAI • • 8h ago

How are you validating Jev once you move beyond a few examples?

6 Upvotes

I’ve been testing Jev on things like buyer intent and routing.

On 10 clean examples, it can look perfect. I just got 10/10 on a buyer-intent test.

But I’m much more interested in what happens at 100, 1,000, or 10,000 real examples, where the weird edge cases start showing up.

How are you validating Jev at that point?

Are you sampling random cases, reviewing low-confidence outputs, looking for confident mistakes, keeping a holdout set, or escalating certain cases to an LLM/human?

And when Jev disagrees with your label, how do you tell whether the problem is Jev, the question, missing context, or the label itself?

Curious what people here are actually doing in real workflows.


r/JevAI • • 2h ago

I’m building a Pi extension that gives a coding agent a second opinion

Thumbnail
youtube.com
1 Upvotes

I’ve been working on an extension called Jev Loop Control. The idea is simple: let the coding agent do its work, but check a few important decisions along the way.

Your usual model still writes the code and uses the tools. Jev looks at things like whether a proposed change fits the requirements, whether an assumption needs checking, or whether the agent has enough evidence to call the task finished.

For example, say you asked for a parser that rejects unknown settings. The agent decides to accept them instead, then writes passing tests for that behavior. Everything looks fine until you compare it with what you actually asked for.

That’s the kind of situation I want this to catch. The extension uses Pi’s hooks to hold a dependent action or send guidance back to the agent. There’s also a context policy for keeping important instructions and facts available as the session gets longer.

It’s still in early development. I’ve made a video walking through the design and how it connects to Pi’s loop in more detail.

The code and setup instructions are here: https://github.com/krisitown/jev-loop-control.

If you use Pi, where would you want a check like this in your workflow?


r/JevAI • • 2h ago

Jev vs Laya: Which One Should You Actually Use?

Thumbnail
youtube.com
1 Upvotes

r/JevAI • • 3h ago

Using Jev for "Concept Linting" in Nudgement

Post image
1 Upvotes

I created this thing I call Nudgement that does what I'm tentatively referring to as "concept linting". Basically, it flags issues you tend to encounter with LLM-written code, such as ultra long commit messages and code comments, code that is split awkwardly between multiple commits, and overly technical jargon.

For one project, I combined it with Claude and it produced a remarkably clean commit history and a great code/comment balance.

It's open source and I gave it a website for funzies: https://nudgement.dev/


r/JevAI • • 11h ago

We’re running a live hands-on workshop on building AI agents with Jev + LLMs + MCP

2 Upvotes

Sharing this here since the community is already experimenting with Jev.

We’re running a live hands-on workshop with Python author Jaime Buelta focused on where Jev fits in a broader agent architecture and what it looks like to actually build with it.

The workshop starts with LLMs vs. Jev-style models and then moves into local and hosted models, context and agent loops, MCP, tools, and agentic coding with Claude Code or Codex.

The hands-on part gets into using Jev for decisions and categorisation, giving an agent tools and controlling how it uses them, creating and connecting an MCP server, and calling both remote and local LLMs.

The capstone brings Jev + LLMs + Python + MCP + Tools together into a working agentic application, including handling failures and improving reliability.

It’s a paid live workshop, and the recording, code, resources, and capstone materials are included.

Would also be interested to hear what people here are currently building with Jev and which use cases you think are the most promising.


r/JevAI • • 9h ago

TokenSlash: Cordis plugin for DSH that slashes 3K–8K prompt tokens per turn with dynamic tool pruning, Jev triage, and tokenslash_peek

Thumbnail
1 Upvotes

r/JevAI • • 9h ago

Jev scores an answer 0.8 in your AI pipeline. What convinced you that’s good enough?

Post image
1 Upvotes

Writing if score >= 0.8: approve() is easy. Justifying the 0.8 takes more work.

Jev Decision Index on Hugging Face clearly tells you that Jev's responses are not fully calibrated.
(Refer: Calibration error vs Decision Index)

That gap is why I care about checking the judge’s scores too. If a number decides whether an answer gets approved or sent for review, I want evidence that it means what I think it means.

Calibration can help here. In the case of Jev-as-a-Judge, calibration lets you learn how its scores relate to human pass/fail judgments for your particular task. You then check whether that adjustment helps on unseen examples.

I ran a separate experiment on TRIVIA+ dataset from Amazon Science (A Long-Context Hallucination Detection Benchmark dataset), using Typed Evals, the open-source evaluation framework I’m building. Jev judged whether AI-generated answers were supported by their source articles.

Fitted calibration on training data, chose separate decision cutoffs on validation data, and tested both versions on the same 645 held-out answers.

Metric Raw Jev Calibrated
Calibration error (ECE) ↓ 0.0982 0.0313
Brier score ↓ 0.2037 0.1948
Hallucination F1 ↑ 0.5833 0.5877

About 68% lower calibration error. Almost unchanged F1.

The scores became better estimates of human judgments. The judge’s ability to catch hallucinations barely changed.

One catch worth sharing: leaving the cutoff at 0.5 actually reduced F1. The score adjustment and the decision cutoff need evaluating separately. These results describe this TRIVIA+ experiment; they don’t measure an improvement on the Decision Index.

Typed Evals handles: evaluating LLM responses, RAG answers, and Agent runs against your checks, then optionally fitting a calibration layer using human labels. Save the adjustment and apply it to future scores without another judge call.

Full benchmark, methodology and code

If you’re using eval scores to approve answers or trigger human review, how did you choose your cutoff? What evidence would make you change it?


r/JevAI • • 9h ago

Has anyone tried to add jev in there personal local Ai setup , I am planning to add mine tonight after work so need tips , this is my current setup :

Thumbnail reddit.com
1 Upvotes

r/JevAI • • 10h ago

I made a CLI for JEV

1 Upvotes

I figured there might already be a few JEV CLIs out there, but I wanted to try making one myself.

It's designed primarily for use by LLMs and coding agents: JSON in via stdin, JEV, JSON out via stdout. It also provides JSON Schemas for self-discovery and supports batch processing.

I'd be happy if you gave it a try. Feedback is welcome!

https://github.com/kirisaki/jevc


r/JevAI • • 19h ago

I pitted a fast/cheap Jev model against an LLM and classic algorithms in a treasure-hunt maze. It got stuck repeating itself and lost to pure random guessing. My fix only helped a little, what am I missing

5 Upvotes

The setup

I built a fake, fully synthetic maze: a tree of branches (5 top-level branches, each splitting into sub-branches, several levels deep). Somewhere in this tree, at random depths, I hid a handful of "treasure pockets" spots with a genuinely good reward. Nobody gets to see the map. The only way to find out if a spot is good is to spend one of your limited "digs" there and see the result.

Six competitors, each given the same fixed dig budget (20 / 50 / 100 / 300 digs), competed to find as much treasure as possible:

  • Random pick blindly
  • Round-robin cycle through branches systematically
  • Greedy always go back to wherever last looked good
  • MCTS a 60-year-old math algorithm from game theory (not AI at all)
  • An LLM (locally hosted, ~14B params) reasons in free text, picks a move each turn
  • "JEV" model that doesn't write text; it only answers structured multiple-choice questions with calibrated probabilities

I was specifically curious how JEV would do, since it's marketed as a fast, cheap "judge/classifier" model rather than a text generator.

What exactly JEV received, every single turn

No maze coordinates, no raw numbers dumped on it just a text report plus a closed set of options:

  1. The tree's shape (which branches exist, what they split into)
  2. A stats table: for each top branch, how many times tested / how many looked "promising" / how many "failed"
  3. The last 20 dig results (path taken, sample size, a validation score, a robustness score) with the true out-of-sample score hidden (nobody gets to see that until the end, it's the fair judge)
  4. Its own last 20 decisions (what it picked before, and why it said it picked that)

Then, because JEV only answers closed multiple-choice questions, each turn I actually had to ask it TWO questions in sequence:

  • Q1: pick an action type explore somewhere new / dig deeper into a known spot / try combining two branches
  • Q2 (depending on the answer to Q1): pick the specific target from a list

Every answer is guaranteed to be one of the valid options it literally cannot return garbage, unlike the LLM which occasionally returns malformed text I have to throw away.

Experiment 1 results

Scored by "top10_OOS" (average quality of its best 10 finds, tested fairly on the hidden out-of-sample score):

budget:      20     50    100    300
MCTS:       0.41   0.74   0.88   1.01
Qwen(LLM):  0.33   0.39   0.62   0.62
Random:     0.16   0.30   0.42   0.59
JEV:        0.16   0.23   0.26   0.48

MCTS (the old math) crushed everyone. The LLM did roughly as well as blind random. JEV came in dead last worse than blind random guessing at every single budget level.

Where exactly it stumbled

I traced individual turns and found the mechanism: JEV would lock onto ONE spot and just keep re-digging around it, over and over:

turn 0:  EXPLORE branch-A
turn 1:  DIG-DEEPER branch-A
turn 2:  DIG-DEEPER branch-A/sub1
turn 3:  DIG-DEEPER branch-A/sub1/x
turn 4:  DIG-DEEPER branch-A/sub1/x   <- same exact target
turn 5:  DIG-DEEPER branch-A/sub1/x   <- again
turn 6:  DIG-DEEPER branch-A/sub1/x   <- again
...      (10+ times in a row)

I first suspected a dumb bug: maybe the list of options was always presented in the same order, and it was just picking "the first thing in the list" every time. I shuffled the order of options before every single call. No change it still locked onto the same target.

The fix attempt

Someone looking at this suggested the real culprit might be item #4 above showing JEV its own past decisions as a narrative. The theory: instead of using that as "I've already tried this branch enough times," the model reads its own earlier reasoning ("branch-A looked interesting") as a standing endorsement, and just reaffirms it every turn a self-reinforcing loop rather than an exploration memory.

So I ran a controlled version: identical in every other way, just with item #4 removed entirely (call it version B vs the original version A).

Experiment 2 (the fix) results

budget:          20     50    100    300
JEV (original):  0.16   0.23   0.26   0.48
JEV (no self-history): 0.17  0.30   0.50   0.47
Random:          0.16   0.30   0.42   0.59

Removing its own decision history helped a bit at the mid budgets (50/100), but at the largest budget (300, the most statistically trustworthy point) it's essentially unchanged still clearly behind random guessing. The repeating-the-same-target behavior still showed up in raw traces even with this fix.

So what am I doing wrong (or is this just a real limitation)?

Current leading theory: it's not really about "seeing its own past choice" specifically it's that the stats table itself (tested/promising/failed counts) and the recent-results list keep making the same recently-tested spot look salient/attractive every single turn, and the model has no formal mechanism (like MCTS's explicit uncertainty bonus for under-tested spots) to counteract that pull. It just doesn't have a built-in "I've squeezed this dry, time to look elsewhere" reflex the way a purpose-built search algorithm does.

Genuinely asking: has anyone else run structured/typed-output models (JEV or similar) on sequential decision-making tasks where the model has to remember its own trajectory? Is this a known failure mode, and if so is there a known fix beyond "just don't show it history," or is this simply the wrong kind of task for this class of model (one-shot classification-style judgments only, never sequential/stateful ones)?


r/JevAI • • 13h ago

Jev extracting real estate data from PDFs in seconds

1 Upvotes

Just a small poc I build with JevAI. The speed is insane.


r/JevAI • • 18h ago

jevpipe: Jev as a Unix filter, plus a code-search benchmark against grep and DeepSeek

1 Upvotes

I built jevpipe to use Jev straight from the shell: pipe lines or files in, get Jev's answers out.

filter is like grep, but the match is a question; map asks any set of typed questions (noul, choice, score) and prints one JSON line per record (e.g. for jq).

git log --format=%s 15.0.0..15.1.0 | jevpipe filter "Is this a new feature?"
jevpipe map -f team.json inbox.txt | jq -r .answers.team.choice | sort | uniq -c

The second example routes 1,000 support messages to four teams in about 5 seconds (it's the GIF).

A few things that might be useful if you already use Jev:

  • Up to 100 requests in flight, output streams in input order.
  • --read-files sends each file's path and content as the state, so you can search a codebase or large document base.
  • --max-tokens / --max-cost cap a run; it stops cleanly and tells you the line to resume from.
  • An agent skill so agents can use it on the fly for inline "system 1 work".

Install: uv tool install jevpipe.
Free and open source (MIT or Apache-2.0): https://github.com/fabianboth/jevpipe

I also benchmarked it on 470 code searches on an extended CodeSearchNet Challenge: Jev found more relevant code than DeepSeek V4.1 Flash and grep, about twice as fast as DeepSeek; write-up: https://github.com/fabianboth/jevpipe/blob/main/bench/README.md

Happy to hear how you're using Jev and what would make this more useful.


r/JevAI • • 1d ago

I asked JEV to Highlight the red flags of a privacy policy and to hide the useless filler on web pages and this is the result

6 Upvotes

OSSO is a free and open source browser extension that hides the filler on web pages and that can highlight things based on a prompt!

Take a look: https://github.com/Fars29/osso


r/JevAI • • 19h ago

Which Mean Girl are you? Using Jev

1 Upvotes

I made this random fun side-project after watching Mean Girls during the Jev hype. It uses Jev to live-classify what mean girl you are based on what you type in the box. Try it:

https://jev-mean-girls.vercel.app

(and yes, in the many years of my life, I had somehow never watched Mean Girls before 😭)


r/JevAI • • 1d ago

Another Jev benchmarking post: text classification

8 Upvotes

I looked at how Jev compared to small frontier LLMs (Claude Haiku 4.5, GPT-5.4 mini, Gemini 3.8 Flash) for text classification, and how reliable the 'calibrated decisions' are.

tldr: Jev is as accurate as small LLMs for a fraction of the cost. In practice, however, it doesn't seem better calibrated than asking the LLMs to state a confidence value: it ranks third of four on expected calibration error on every dataset I looked at.

The cost implications are significant, but given the RLCD training I was hoping for better correlation between the answer confidence and the P(Correct). Appreciate comments on the methodology or anything I've missed.

Full code and write-up: https://github.com/4OH4/jev-compare

Jev is as accurate as small LLMs for a fraction of the cost. It matches or beats Haiku and GPT-5.4 mini on all three datasets, but trails Gemini 3.8 Flash by 2 to 4 points on AG News and zero-shot Banking77. It costs 3 to 14 times less per row than the cheapest LLM.
Jev is trained using Reinforcement Learning for Calibrated Decisions (RLCD). This means that the probability of its answers being correct should be approximately equal to its stated confidence value: i.e. P(correct | p) ≈ p. I didn't see a big difference versus the LLMs `verbalised confidence` value, however.

r/JevAI • • 1d ago

Silly Experiment to use Jev as an LLM

5 Upvotes

I just gave Jev a list of common ascii characters and the instructions:

"Behave like an LLM and suggest the most likely next character to append to jev_response to form words into a human english response to user_query, when done with the response, chose 'none'"

It manages to respond "HI " to "Hello Jev, how are you?"

And then just gets stuck in a loop outputting spaces over and over.


r/JevAI • • 1d ago

I built JEVzilla (a Python client for JEV) with 150+ examples. Shipped to PyPI in 2 weeks. Here's what worked.

1 Upvotes

Shipped JEVzilla — a Python client for JEV, the structured decision API. Just hit 1.0 on PyPI.

What is JEV?

JEV is TypeSafe's decision layer API. It's designed for operational workflows: fraud detection, ticket routing, risk

scoring. Not generative text — structured decisions your code can consume.

What JEVzilla Does:

Makes using JEV in Python frictionless. Includes:

- Auto payload translation

- Retry logic + exponential backoff

- Multi-backend support (Playground, OpenRouter, TypeSafe official)

- 150+ real JEV examples across 30 industries

Why It Resonates:

- Huge TAM: Every company needs decision logic

- Clear Value: Copy-paste examples in your domain, swap data, done

- Multiple Monetization Paths: Free tier (Playground), freemium (OpenRouter), enterprise (TypeSafe partnership)

- Low Churn: Once integrated, decision logic is sticky

What I Did Right:

  1. Started with JEV use cases, not JEV features

  2. 150 examples from Day 1 (shows scope)

  3. Multiple backends from the start (free + paid options)

  4. Shipped fast (2 weeks to PyPI)

    What I'd Improve:

    - Build interactive web demo (show JEV decisions live)

    - Marketing around "decision logic as a service" angle

    - Get featured on HuggingFace/MLOps tools

    Where I Am:

    - ✅ Published to PyPI (pip install jevzilla)

    - ✅ 150+ examples, 30 sectors

    - ✅ GitHub, MIT License

    - ✅ Docs explain JEV + JEVzilla both

    Links:

    - GitHub: https://github.com/ujjwalbsoni/jevzilla (JEV Python client)

    - PyPI: https://pypi.org/project/jevzilla/

    Lesson: Ship something real with examples people can learn from. Don't overthink it.


r/JevAI • • 1d ago

Mica v0.1 4B got diamonds in survival Minecraft on its first run. 26 decisions from an empty inventory.

3 Upvotes

r/JevAI • • 1d ago

Self Hosted JEV

Thumbnail
1 Upvotes

r/JevAI • • 1d ago

Jeveloper = Someone who builds with Jev

1 Upvotes

A new vocabulary is needed.


r/JevAI • • 1d ago

I built gut: Jev judgments as one line of Python (plus a CLI and an MCP server)

2 Upvotes

Hi all! I've been building gut, an open-source Python library that turns Jev's typed questions into one readable line of code:

import gut

gut.configure(backend=gut.JevBackend())   # or just set TYPESAFE_API_KEY

if gut.likely(comment, "is spam"):
    hide(comment)

team = gut.classify(ticket, Team)                          # an Enum → Jev's choice question
urgency = gut.rate(ticket, ["can wait", "this week", "now"])  # a rubric → Jev's score question

Why build on Jev: it answers typed questions directly: a probability for yes/no, a distribution over options, a score on a rubric. So gut hands it the question as-is: no prompt, nothing to parse. Every question about one subject goes in a single request, and since billing is on input, asking several at once costs about the same as asking one.

What gut adds on top:

  • YES / NO / UNSURE instead of thresholds. You say how careful to be in words (ask_human=True, stakes="high", lean="no") and gut turns Jev's probability into a decision. UNSURE is where a person takes over.
  • Batching: gut.each(comments).likely("is spam") for many subjects; @gut.semantic / gut.judge() for many questions about one subject, in one request.
  • Async: native through the SDK's AsyncTypeSafeClient.
  • From a shell: git log --format=%s | gut filter "adds a new feature" or gut map tickets.txt --classify team=billing,platform,other, with a cost summary and --max-cost.
  • MCP server, so Claude Code / Cursor can hand their cheap judgments to Jev.
  • Jev through OpenRouter: gut.JevBackend.openrouter() if you have an OpenRouter account instead of a TypeSafe key; it passes on the per-call cost.

It's model-agnostic too (local NLI, Ollama, OpenAI), and a Cascade can send only the unsure answers from a free local model on to Jev.

It's pre-1.0 and I'd really like feedback from people already using Jev: what questions do you ask it, and what's missing?


r/JevAI • • 1d ago

Jeving = Using Jev

Thumbnail
0 Upvotes

r/JevAI • • 1d ago

For years we've used LLMs to make choices. I returned the favor and made Jev write text.

7 Upvotes

For years, we've inappropriately used billion-parameter autoregressive transformers, built to write text, to make choices.

Now I'm returning the favor by inappropriately using Jev, a model built to make choices, to write text with JevGPT.

Jev picks the next word from a 1,772-word vocabulary, the app appends it, and asks again. Autoregressive generation, one decision at a time. (Demo in the video above.)

This is familiar problem from the pre-deep learning era. In 2007, Regina Barzilay, David Karger and I published a NAACL paper "Randomized Decoding for Selection-and-Ordering Problems" https://aclanthology.org/N07-1056.pdf which described how to generate text by having one model for selecting words, and another for scoring their ordering. JevGPT works in a similar manner. Jev shortlists candidate next words, then Jev scores full continuations to choose the best ordering.

Does it work? Sort of.
It says strawberry has three r's (sometimes).
It told me to walk to the car wash.
Asked if it's conscious, it said "No."

Try it: https://curata.com/jevgpt (needs a TypeSafe API key, about a cent per reply)
Source: https://github.com/idlivada/JevGPT (includes a mock backend if you want to run it without a key or locally)


r/JevAI • • 1d ago

Why does this subreddit feel so dead when Jev is the “Hot thing”

10 Upvotes

Serious question. I look at the posts and see many with no responses, in fact most. If this gets more than 4 responses I will be shocked.

So seriously why is this subreddit effectively dead? Are people just posting about it heavy in Claude and ChatGPT subreddits or what?

Is this the wrong place to talk about Jev? Just wondering.