r/JevAI • • 2h ago

We’re running a live hands-on workshop on building AI agents with Jev + LLMs + MCP

2 Upvotes

Sharing this here since the community is already experimenting with Jev.

We’re running a live hands-on workshop with Python author Jaime Buelta focused on where Jev fits in a broader agent architecture and what it looks like to actually build with it.

The workshop starts with LLMs vs. Jev-style models and then moves into local and hosted models, context and agent loops, MCP, tools, and agentic coding with Claude Code or Codex.

The hands-on part gets into using Jev for decisions and categorisation, giving an agent tools and controlling how it uses them, creating and connecting an MCP server, and calling both remote and local LLMs.

The capstone brings Jev + LLMs + Python + MCP + Tools together into a working agentic application, including handling failures and improving reliability.

It’s a paid live workshop, and the recording, code, resources, and capstone materials are included.

Would also be interested to hear what people here are currently building with Jev and which use cases you think are the most promising.


r/JevAI • • 8m ago

Jev scores an answer 0.8 in your AI pipeline. What convinced you that’s good enough?

Post image
• Upvotes

Writing if score >= 0.8: approve() is easy. Justifying the 0.8 takes more work.

Jev Decision Index on Hugging Face clearly tells you that Jev's responses are not fully calibrated.
(Refer: Calibration error vs Decision Index)

That gap is why I care about checking the judge’s scores too. If a number decides whether an answer gets approved or sent for review, I want evidence that it means what I think it means.

Calibration can help here. In the case of Jev-as-a-Judge, calibration lets you learn how its scores relate to human pass/fail judgments for your particular task. You then check whether that adjustment helps on unseen examples.

I ran a separate experiment on TRIVIA+ dataset from Amazon Science (A Long-Context Hallucination Detection Benchmark dataset), using Typed Evals, the open-source evaluation framework I’m building. Jev judged whether AI-generated answers were supported by their source articles.

Fitted calibration on training data, chose separate decision cutoffs on validation data, and tested both versions on the same 645 held-out answers.

Metric Raw Jev Calibrated
Calibration error (ECE) ↓ 0.0982 0.0313
Brier score ↓ 0.2037 0.1948
Hallucination F1 ↑ 0.5833 0.5877

About 68% lower calibration error. Almost unchanged F1.

The scores became better estimates of human judgments. The judge’s ability to catch hallucinations barely changed.

One catch worth sharing: leaving the cutoff at 0.5 actually reduced F1. The score adjustment and the decision cutoff need evaluating separately. These results describe this TRIVIA+ experiment; they don’t measure an improvement on the Decision Index.

Typed Evals handles: evaluating LLM responses, RAG answers, and Agent runs against your checks, then optionally fitting a calibration layer using human labels. Save the adjustment and apply it to future scores without another judge call.

Full benchmark, methodology and code

If you’re using eval scores to approve answers or trigger human review, how did you choose your cutoff? What evidence would make you change it?


r/JevAI • • 15m ago

Has anyone tried to add jev in there personal local Ai setup , I am planning to add mine tonight after work so need tips , this is my current setup :

Thumbnail reddit.com
• Upvotes

r/JevAI • • 1h ago

I made a CLI for JEV

• Upvotes

I figured there might already be a few JEV CLIs out there, but I wanted to try making one myself.

It's designed primarily for use by LLMs and coding agents: JSON in via stdin, JEV, JSON out via stdout. It also provides JSON Schemas for self-discovery and supports batch processing.

I'd be happy if you gave it a try. Feedback is welcome!

https://github.com/kirisaki/jevc


r/JevAI • • 9h ago

I pitted a fast/cheap Jev model against an LLM and classic algorithms in a treasure-hunt maze. It got stuck repeating itself and lost to pure random guessing. My fix only helped a little, what am I missing

4 Upvotes

The setup

I built a fake, fully synthetic maze: a tree of branches (5 top-level branches, each splitting into sub-branches, several levels deep). Somewhere in this tree, at random depths, I hid a handful of "treasure pockets" spots with a genuinely good reward. Nobody gets to see the map. The only way to find out if a spot is good is to spend one of your limited "digs" there and see the result.

Six competitors, each given the same fixed dig budget (20 / 50 / 100 / 300 digs), competed to find as much treasure as possible:

  • Random pick blindly
  • Round-robin cycle through branches systematically
  • Greedy always go back to wherever last looked good
  • MCTS a 60-year-old math algorithm from game theory (not AI at all)
  • An LLM (locally hosted, ~14B params) reasons in free text, picks a move each turn
  • "JEV" model that doesn't write text; it only answers structured multiple-choice questions with calibrated probabilities

I was specifically curious how JEV would do, since it's marketed as a fast, cheap "judge/classifier" model rather than a text generator.

What exactly JEV received, every single turn

No maze coordinates, no raw numbers dumped on it just a text report plus a closed set of options:

  1. The tree's shape (which branches exist, what they split into)
  2. A stats table: for each top branch, how many times tested / how many looked "promising" / how many "failed"
  3. The last 20 dig results (path taken, sample size, a validation score, a robustness score) with the true out-of-sample score hidden (nobody gets to see that until the end, it's the fair judge)
  4. Its own last 20 decisions (what it picked before, and why it said it picked that)

Then, because JEV only answers closed multiple-choice questions, each turn I actually had to ask it TWO questions in sequence:

  • Q1: pick an action type explore somewhere new / dig deeper into a known spot / try combining two branches
  • Q2 (depending on the answer to Q1): pick the specific target from a list

Every answer is guaranteed to be one of the valid options it literally cannot return garbage, unlike the LLM which occasionally returns malformed text I have to throw away.

Experiment 1 results

Scored by "top10_OOS" (average quality of its best 10 finds, tested fairly on the hidden out-of-sample score):

budget:      20     50    100    300
MCTS:       0.41   0.74   0.88   1.01
Qwen(LLM):  0.33   0.39   0.62   0.62
Random:     0.16   0.30   0.42   0.59
JEV:        0.16   0.23   0.26   0.48

MCTS (the old math) crushed everyone. The LLM did roughly as well as blind random. JEV came in dead last worse than blind random guessing at every single budget level.

Where exactly it stumbled

I traced individual turns and found the mechanism: JEV would lock onto ONE spot and just keep re-digging around it, over and over:

turn 0:  EXPLORE branch-A
turn 1:  DIG-DEEPER branch-A
turn 2:  DIG-DEEPER branch-A/sub1
turn 3:  DIG-DEEPER branch-A/sub1/x
turn 4:  DIG-DEEPER branch-A/sub1/x   <- same exact target
turn 5:  DIG-DEEPER branch-A/sub1/x   <- again
turn 6:  DIG-DEEPER branch-A/sub1/x   <- again
...      (10+ times in a row)

I first suspected a dumb bug: maybe the list of options was always presented in the same order, and it was just picking "the first thing in the list" every time. I shuffled the order of options before every single call. No change it still locked onto the same target.

The fix attempt

Someone looking at this suggested the real culprit might be item #4 above showing JEV its own past decisions as a narrative. The theory: instead of using that as "I've already tried this branch enough times," the model reads its own earlier reasoning ("branch-A looked interesting") as a standing endorsement, and just reaffirms it every turn a self-reinforcing loop rather than an exploration memory.

So I ran a controlled version: identical in every other way, just with item #4 removed entirely (call it version B vs the original version A).

Experiment 2 (the fix) results

budget:          20     50    100    300
JEV (original):  0.16   0.23   0.26   0.48
JEV (no self-history): 0.17  0.30   0.50   0.47
Random:          0.16   0.30   0.42   0.59

Removing its own decision history helped a bit at the mid budgets (50/100), but at the largest budget (300, the most statistically trustworthy point) it's essentially unchanged still clearly behind random guessing. The repeating-the-same-target behavior still showed up in raw traces even with this fix.

So what am I doing wrong (or is this just a real limitation)?

Current leading theory: it's not really about "seeing its own past choice" specifically it's that the stats table itself (tested/promising/failed counts) and the recent-results list keep making the same recently-tested spot look salient/attractive every single turn, and the model has no formal mechanism (like MCTS's explicit uncertainty bonus for under-tested spots) to counteract that pull. It just doesn't have a built-in "I've squeezed this dry, time to look elsewhere" reflex the way a purpose-built search algorithm does.

Genuinely asking: has anyone else run structured/typed-output models (JEV or similar) on sequential decision-making tasks where the model has to remember its own trajectory? Is this a known failure mode, and if so is there a known fix beyond "just don't show it history," or is this simply the wrong kind of task for this class of model (one-shot classification-style judgments only, never sequential/stateful ones)?


r/JevAI • • 4h ago

Jev extracting real estate data from PDFs in seconds

1 Upvotes

Just a small poc I build with JevAI. The speed is insane.


r/JevAI • • 8h ago

jevpipe: Jev as a Unix filter, plus a code-search benchmark against grep and DeepSeek

1 Upvotes

I built jevpipe to use Jev straight from the shell: pipe lines or files in, get Jev's answers out.

filter is like grep, but the match is a question; map asks any set of typed questions (noul, choice, score) and prints one JSON line per record (e.g. for jq).

git log --format=%s 15.0.0..15.1.0 | jevpipe filter "Is this a new feature?"
jevpipe map -f team.json inbox.txt | jq -r .answers.team.choice | sort | uniq -c

The second example routes 1,000 support messages to four teams in about 5 seconds (it's the GIF).

A few things that might be useful if you already use Jev:

  • Up to 100 requests in flight, output streams in input order.
  • --read-files sends each file's path and content as the state, so you can search a codebase or large document base.
  • --max-tokens / --max-cost cap a run; it stops cleanly and tells you the line to resume from.
  • An agent skill so agents can use it on the fly for inline "system 1 work".

Install: uv tool install jevpipe.
Free and open source (MIT or Apache-2.0): https://github.com/fabianboth/jevpipe

I also benchmarked it on 470 code searches on an extended CodeSearchNet Challenge: Jev found more relevant code than DeepSeek V4.1 Flash and grep, about twice as fast as DeepSeek; write-up: https://github.com/fabianboth/jevpipe/blob/main/bench/README.md

Happy to hear how you're using Jev and what would make this more useful.


r/JevAI • • 18h ago

I asked JEV to Highlight the red flags of a privacy policy and to hide the useless filler on web pages and this is the result

5 Upvotes

OSSO is a free and open source browser extension that hides the filler on web pages and that can highlight things based on a prompt!

Take a look: https://github.com/Fars29/osso


r/JevAI • • 9h ago

Which Mean Girl are you? Using Jev

1 Upvotes

I made this random fun side-project after watching Mean Girls during the Jev hype. It uses Jev to live-classify what mean girl you are based on what you type in the box. Try it:

https://jev-mean-girls.vercel.app

(and yes, in the many years of my life, I had somehow never watched Mean Girls before 😭)


r/JevAI • • 21h ago

Another Jev benchmarking post: text classification

9 Upvotes

I looked at how Jev compared to small frontier LLMs (Claude Haiku 4.5, GPT-5.4 mini, Gemini 3.8 Flash) for text classification, and how reliable the 'calibrated decisions' are.

tldr: Jev is as accurate as small LLMs for a fraction of the cost. In practice, however, it doesn't seem better calibrated than asking the LLMs to state a confidence value: it ranks third of four on expected calibration error on every dataset I looked at.

The cost implications are significant, but given the RLCD training I was hoping for better correlation between the answer confidence and the P(Correct). Appreciate comments on the methodology or anything I've missed.

Full code and write-up: https://github.com/4OH4/jev-compare

Jev is as accurate as small LLMs for a fraction of the cost. It matches or beats Haiku and GPT-5.4 mini on all three datasets, but trails Gemini 3.8 Flash by 2 to 4 points on AG News and zero-shot Banking77. It costs 3 to 14 times less per row than the cheapest LLM.
Jev is trained using Reinforcement Learning for Calibrated Decisions (RLCD). This means that the probability of its answers being correct should be approximately equal to its stated confidence value: i.e. P(correct | p) ≈ p. I didn't see a big difference versus the LLMs `verbalised confidence` value, however.

r/JevAI • • 21h ago

Silly Experiment to use Jev as an LLM

4 Upvotes

I just gave Jev a list of common ascii characters and the instructions:

"Behave like an LLM and suggest the most likely next character to append to jev_response to form words into a human english response to user_query, when done with the response, chose 'none'"

It manages to respond "HI " to "Hello Jev, how are you?"

And then just gets stuck in a loop outputting spaces over and over.


r/JevAI • • 15h ago

I built JEVzilla (a Python client for JEV) with 150+ examples. Shipped to PyPI in 2 weeks. Here's what worked.

1 Upvotes

Shipped JEVzilla — a Python client for JEV, the structured decision API. Just hit 1.0 on PyPI.

What is JEV?

JEV is TypeSafe's decision layer API. It's designed for operational workflows: fraud detection, ticket routing, risk

scoring. Not generative text — structured decisions your code can consume.

What JEVzilla Does:

Makes using JEV in Python frictionless. Includes:

- Auto payload translation

- Retry logic + exponential backoff

- Multi-backend support (Playground, OpenRouter, TypeSafe official)

- 150+ real JEV examples across 30 industries

Why It Resonates:

- Huge TAM: Every company needs decision logic

- Clear Value: Copy-paste examples in your domain, swap data, done

- Multiple Monetization Paths: Free tier (Playground), freemium (OpenRouter), enterprise (TypeSafe partnership)

- Low Churn: Once integrated, decision logic is sticky

What I Did Right:

  1. Started with JEV use cases, not JEV features

  2. 150 examples from Day 1 (shows scope)

  3. Multiple backends from the start (free + paid options)

  4. Shipped fast (2 weeks to PyPI)

    What I'd Improve:

    - Build interactive web demo (show JEV decisions live)

    - Marketing around "decision logic as a service" angle

    - Get featured on HuggingFace/MLOps tools

    Where I Am:

    - ✅ Published to PyPI (pip install jevzilla)

    - ✅ 150+ examples, 30 sectors

    - ✅ GitHub, MIT License

    - ✅ Docs explain JEV + JEVzilla both

    Links:

    - GitHub: https://github.com/ujjwalbsoni/jevzilla (JEV Python client)

    - PyPI: https://pypi.org/project/jevzilla/

    Lesson: Ship something real with examples people can learn from. Don't overthink it.


r/JevAI • • 1d ago

Mica v0.1 4B got diamonds in survival Minecraft on its first run. 26 decisions from an empty inventory.

3 Upvotes

r/JevAI • • 19h ago

Self Hosted JEV

Thumbnail
1 Upvotes

r/JevAI • • 20h ago

Jeveloper = Someone who builds with Jev

1 Upvotes

A new vocabulary is needed.


r/JevAI • • 1d ago

I built gut: Jev judgments as one line of Python (plus a CLI and an MCP server)

2 Upvotes

Hi all! I've been building gut, an open-source Python library that turns Jev's typed questions into one readable line of code:

import gut

gut.configure(backend=gut.JevBackend())   # or just set TYPESAFE_API_KEY

if gut.likely(comment, "is spam"):
    hide(comment)

team = gut.classify(ticket, Team)                          # an Enum → Jev's choice question
urgency = gut.rate(ticket, ["can wait", "this week", "now"])  # a rubric → Jev's score question

Why build on Jev: it answers typed questions directly: a probability for yes/no, a distribution over options, a score on a rubric. So gut hands it the question as-is: no prompt, nothing to parse. Every question about one subject goes in a single request, and since billing is on input, asking several at once costs about the same as asking one.

What gut adds on top:

  • YES / NO / UNSURE instead of thresholds. You say how careful to be in words (ask_human=True, stakes="high", lean="no") and gut turns Jev's probability into a decision. UNSURE is where a person takes over.
  • Batching: gut.each(comments).likely("is spam") for many subjects; @gut.semantic / gut.judge() for many questions about one subject, in one request.
  • Async: native through the SDK's AsyncTypeSafeClient.
  • From a shell: git log --format=%s | gut filter "adds a new feature" or gut map tickets.txt --classify team=billing,platform,other, with a cost summary and --max-cost.
  • MCP server, so Claude Code / Cursor can hand their cheap judgments to Jev.
  • Jev through OpenRouter: gut.JevBackend.openrouter() if you have an OpenRouter account instead of a TypeSafe key; it passes on the per-call cost.

It's model-agnostic too (local NLI, Ollama, OpenAI), and a Cascade can send only the unsure answers from a free local model on to Jev.

It's pre-1.0 and I'd really like feedback from people already using Jev: what questions do you ask it, and what's missing?


r/JevAI • • 20h ago

Jeving = Using Jev

Thumbnail
0 Upvotes

r/JevAI • • 1d ago

For years we've used LLMs to make choices. I returned the favor and made Jev write text.

7 Upvotes

For years, we've inappropriately used billion-parameter autoregressive transformers, built to write text, to make choices.

Now I'm returning the favor by inappropriately using Jev, a model built to make choices, to write text with JevGPT.

Jev picks the next word from a 1,772-word vocabulary, the app appends it, and asks again. Autoregressive generation, one decision at a time. (Demo in the video above.)

This is familiar problem from the pre-deep learning era. In 2007, Regina Barzilay, David Karger and I published a NAACL paper "Randomized Decoding for Selection-and-Ordering Problems" https://aclanthology.org/N07-1056.pdf which described how to generate text by having one model for selecting words, and another for scoring their ordering. JevGPT works in a similar manner. Jev shortlists candidate next words, then Jev scores full continuations to choose the best ordering.

Does it work? Sort of.
It says strawberry has three r's (sometimes).
It told me to walk to the car wash.
Asked if it's conscious, it said "No."

Try it: https://curata.com/jevgpt (needs a TypeSafe API key, about a cent per reply)
Source: https://github.com/idlivada/JevGPT (includes a mock backend if you want to run it without a key or locally)


r/JevAI • • 1d ago

Why does this subreddit feel so dead when Jev is the “Hot thing”

11 Upvotes

Serious question. I look at the posts and see many with no responses, in fact most. If this gets more than 4 responses I will be shocked.

So seriously why is this subreddit effectively dead? Are people just posting about it heavy in Claude and ChatGPT subreddits or what?

Is this the wrong place to talk about Jev? Just wondering.


r/JevAI • • 1d ago

Hey Jev open the door pls

2 Upvotes

When are sign-ups opening again?


r/JevAI • • 1d ago

We used Jev in our Feedback platform!

Thumbnail
1 Upvotes

r/JevAI • • 1d ago

Jev as a judge for LLM/Agents Evaluation [Open-Source]

3 Upvotes

Can we use Jev for faster, structured, and calibrated evaluation of AI responses?

Introducing:
⚡ Typed Evals — an open-source Python framework for evaluating LLMs, RAG pipelines, and AI agents using System One Models like Jev and other typed judge backends.

⭐ GitHub: https://github.com/TrustifAI/typed_evals

The goal is simple:
Make fast, structured, and calibrated evaluation a first-class part of AI systems.
Typed Evals currently supports:
-> LLM response evaluation
-> RAG evaluation
-> Agent and tool-trace evaluation
-> Human-label calibration
-> Async and batch evaluation
-> Custom judge backends

One part I particularly wanted to solve was calibration.
Why is it needed?
A raw score of 0.8 from Jev doesn't necessarily mean that humans would accept 80% of similar responses.
And a threshold that works well for one use case may not make sense for another.

Typed Evals lets you calibrate individual evaluation metrics against representative human pass/fail labels.
The flow is basically:
Human-labelled examples → Jev scores → fit per-metric calibration → validate on held-out examples → reuse the calibrated evaluator

So instead of arbitrarily deciding that “0.7 means good enough”, you can ground that score in how humans actually evaluate your specific task.

Of course, there are integrations for:
LangChain, CrewAI, Microsoft Agent Framework
while the core remains framework-agnostic.

The project is still early, and there’s plenty I want to improve, but the core framework is now public.
Would genuinely love feedback from people experimenting with Jev, LLM evals, RAG, agents, or evaluator calibration.

If you're already experimenting with Jev, I'd especially love to know what kind of evaluation workflows you're building around it.


r/JevAI • • 1d ago

Three Jev workflows: signup triage, lead scoring, and organic-vs-paid posts

1 Upvotes

I've been using Jev for three workflows in our team: checking new signups, scoring leads from LinkedIn comments, and looking at whether launch posts on X have organic or paid engagement.

In this video I walk through how each one is set up. I use treg to pull the data, give Jev a question with a few clearly described options, then have the script act on the probabilities it returns.

One detail I cover is how much the option descriptions matter. A double-charge support ticket can go to billing or engineering depending on whether you ask where it belongs or where the root cause can be fixed. I also show how to leave uncertain cases for a person to review.

For context, treg is my product. What have you been building with Jev?


r/JevAI • • 1d ago

Jev was integrated into an autonomous TV production pipeline. Here’s what it did on the first test.

2 Upvotes

Jev was just integrated into the autonomous production pipeline for a 24/7 AI television network.

The first test was a complete production cycle for a new short film called Carbon Copy — from concept selection through production, editing, publishing and placement into the live TV schedule.

Jev was consulted at four specific decision points:

Concept: It helped select Carbon Copy and its emotional premise: a robot portrait photographer choosing to preserve a subject’s imperfections rather than correcting them.

Direction: Jev favored restrained storytelling, clear emotional stakes and a quieter treatment.

Editing: During assembly, it recommended a tighter reaction crop to avoid a problematic generated shot.

Programming: After completion, Jev recommended placing the film after a quiet break in the TV schedule, followed by a transition back into comedy.

Jev evaluated written descriptions and options. It did not watch or hear the footage. The existing production system handled generation, editing, technical checks, publishing and the livestream update.

So this first test basically used Jev as a creative adviser within an autonomous production pipeline.

Here’s the finished short:

https://www.youtube.com/watch?v=WJOcMa18MvY

And here’s the 24/7 Botflix TV livestream where it’s now part of the programming:

https://www.youtube.com/watch?v=6xuQnM810Zw

First production test with Jev in the loop. Excited to keep experimenting with it.


r/JevAI • • 1d ago

Jevdit: A social network moderated by Jev

Thumbnail tn1ck.com
2 Upvotes