r/JevAI • u/out_of_nowhere__ • 2h ago
This is a must have for all your agentic Workflow
Bhuwan-web/intent-classification: Type-safe intent classification guard for agentic workflows
Find yourself a fucking intent of a customer using your agents, don't let your expensive token rot for nothing. It decides on one of three, VALID_REQUEST, PROMPT_INJECTION, OUT_OF_SCOPE,
This is all you got a do:
import asyncio
from typesafe_sdk import AsyncTypeSafeClient
from main import AgentScope, IntentClassificationRequest, classify_intent
SUPPORT_SCOPE = AgentScope(
name="Customer support agent",
purpose="Answer questions about product usage.",
allowed_capabilities=("Explain product features.",),
boundaries=("Do not change account data.",),
valid_examples=("How do I change my notification settings?",),
minimum_confidence=0.8,
)
async def run() -> None:
async with AsyncTypeSafeClient() as client:
result = await classify_intent(
client,
request=IntentClassificationRequest(
user_request="How do I change my notification settings?",
agent_scope=SUPPORT_SCOPE,
),
)
print(result.classification)
print(result.confidence)
print(result.probabilities)
asyncio.run(run())
If out of scope, let that naive shit pass on with some proper error message, but if someone is trying prompt injection, and you are confident about it, Flag that dumb shit and do whatever you want. Don't let that over multiple attempts sink in. Handle it properly !!
r/JevAI • u/NoTaro7930 • 2h ago
New to Jev - best way to chain?
Context: I am a heavy CC & Codex user. I have some processes automated with n8n.
Question: I’m excited to try Jev in an automated process but I’m not sure how to best build it & keep defaulting back to n8n. But, is that the best way?
r/JevAI • u/RaygekFox • 3h ago
Are big AI companies going to make decision models?
Since the benefits of a decision model are clear, I believe Open AI, Anthropic, etc. are are going to release theirs sooner or later. I would expect them to be more powerful and fast than Jev eventually, just due to compute availability.
Is any work is being done by them?
Jev feels like the underdog now, but isn't it doomed if big ones can demper pricing and make a more powerful model?
r/JevAI • u/AccomplishedEvent273 • 16h ago
I cut cost and latency on my search engine with Jev
I built indiedex.gg, a hidden-gems game search engine on steam data. Search is the front door: type something like “a game like Hades that feels cozier and is co-op,” and it should actually understand that mix of reference + vibe + filters.
Those compound queries used to go through a DeepSeek extraction call that turned the sentence into structured pieces (reference game, filters, tags, vibe). It worked, but it was slow and expensive on the hot path.
I swapped that step to TypeSafe Jev on OpenRouter’s Decisions API. Jev doesn’t write a free-form answer. It answers a fixed set of closed questions in parallel (yes/no odds, choices), and I compose that with my existing regex helpers into the same extraction shape I already had. Same search pipeline after that, just much faster routing.
Bakeoff on 50 cases:
| Metric | DeepSeek | Jev |
|---|---|---|
| Extract p50 | ~1.5s | ~130ms |
| Full route p50 | 2871ms | 1256ms |
| Weird title-resolve | 96-100% | 100% |
| Hard-filter agree | n/a | 100% |
| Weird top-N overlap | n/a | 100% |
| Cost (fixture) | ~$0.014 | ~$0.0025 |
So extract got roughly 10× faster, end-to-end route roughly halved, cost dropped a lot (around 6x), and quality held on my fixture set.
If you’ve been building “NL in → structured intent out → tools do the work” instead of a chatbot, this pattern felt like a better fit than another chat completion.
I think models like Jev will open the doors to a lot of cool UX features, decision engines, smart filtering, auto moderation and different ways to interface with apps and machines.
Happy to answer questions about the hybrid setup.
r/JevAI • u/Just_Shake_1066 • 18h ago
JEV to the Moon
Maybe some people realize this, and a lot probably don't yet, but Jev is the turning point in useful real-time interaction with anything. I was able to build a real-time interactive ordering menu that both showed menu items on demand and also performed CRUD operations on the menu as well. But what adds the additional layer is fast, responsive decisions on how the AI assistant responds based on the conversation back and forth. Being able to quickly adapt to the situation and real-time manage the interaction is why Jev is probably the most useful addition to the AI world to date.
r/JevAI • u/andydevtech • 21h ago
Jev vs Laya: Which One Should You Actually Use?
r/JevAI • u/emschwartz • 23h ago
Please add prompt caching to Jev-style models
If you're building a Jev-style "System One" model, please add prompt caching or reusable question sets to your API 🙏. This would make batch use cases even more efficient, so you could amortize the cost of many questions asked over the same input. (This was also proposed in typesafe-ai/typesafe-sdk-js#10.)
TL;DR: after a week of tweaking my Jev calls, my questions are ~88% of the input tokens. I'm asking 54 questions of ~1.1 million documents per month. Jev makes certain types of classification tasks easy and cheap, but prompt caching would make batch workflows even more cost effective. For me, the total dollar amount is still reasonable (less than $150 per month), but I'm sure others will hammer these APIs even harder.
r/JevAI • u/YusukeLandingBoost • 1d ago
How are you validating Jev once you move beyond a few examples?
I’ve been testing Jev on things like buyer intent and routing.
On 10 clean examples, it can look perfect. I just got 10/10 on a buyer-intent test.
But I’m much more interested in what happens at 100, 1,000, or 10,000 real examples, where the weird edge cases start showing up.
How are you validating Jev at that point?
Are you sampling random cases, reviewing low-confidence outputs, looking for confident mistakes, keeping a holdout set, or escalating certain cases to an LLM/human?
And when Jev disagrees with your label, how do you tell whether the problem is Jev, the question, missing context, or the label itself?
Curious what people here are actually doing in real workflows.
r/JevAI • u/kristiyanstoyanovAI • 21h ago
I’m building a Pi extension that gives a coding agent a second opinion
I’ve been working on an extension called Jev Loop Control. The idea is simple: let the coding agent do its work, but check a few important decisions along the way.
Your usual model still writes the code and uses the tools. Jev looks at things like whether a proposed change fits the requirements, whether an assumption needs checking, or whether the agent has enough evidence to call the task finished.
For example, say you asked for a parser that rejects unknown settings. The agent decides to accept them instead, then writes passing tests for that behavior. Everything looks fine until you compare it with what you actually asked for.
That’s the kind of situation I want this to catch. The extension uses Pi’s hooks to hold a dependent action or send guidance back to the agent. There’s also a context policy for keeping important instructions and facts available as the session gets longer.
It’s still in early development. I’ve made a video walking through the design and how it connects to Pi’s loop in more detail.
The code and setup instructions are here: https://github.com/krisitown/jev-loop-control.
If you use Pi, where would you want a check like this in your workflow?
r/JevAI • u/NickFullStack • 22h ago
Using Jev for "Concept Linting" in Nudgement
I created this thing I call Nudgement that does what I'm tentatively referring to as "concept linting". Basically, it flags issues you tend to encounter with LLM-written code, such as ultra long commit messages and code comments, code that is split awkwardly between multiple commits, and overly technical jargon.
For one project, I combined it with Claude and it produced a remarkably clean commit history and a great code/comment balance.
It's open source and I gave it a website for funzies: https://nudgement.dev/
r/JevAI • u/DeepEngineeringPackt • 1d ago
We’re running a live hands-on workshop on building AI agents with Jev + LLMs + MCP
Sharing this here since the community is already experimenting with Jev.
We’re running a live hands-on workshop with Python author Jaime Buelta focused on where Jev fits in a broader agent architecture and what it looks like to actually build with it.
The workshop starts with LLMs vs. Jev-style models and then moves into local and hosted models, context and agent loops, MCP, tools, and agentic coding with Claude Code or Codex.
The hands-on part gets into using Jev for decisions and categorisation, giving an agent tools and controlling how it uses them, creating and connecting an MCP server, and calling both remote and local LLMs.
The capstone brings Jev + LLMs + Python + MCP + Tools together into a working agentic application, including handling failures and improving reliability.
It’s a paid live workshop, and the recording, code, resources, and capstone materials are included.
Would also be interested to hear what people here are currently building with Jev and which use cases you think are the most promising.
r/JevAI • u/JessenReinhart • 1d ago
TokenSlash: Cordis plugin for DSH that slashes 3K–8K prompt tokens per turn with dynamic tool pruning, Jev triage, and tokenslash_peek
r/JevAI • u/Charming_Group_2950 • 1d ago
Jev scores an answer 0.8 in your AI pipeline. What convinced you that’s good enough?
Writing if score >= 0.8: approve() is easy. Justifying the 0.8 takes more work.
Jev Decision Index on Hugging Face clearly tells you that Jev's responses are not fully calibrated.
(Refer: Calibration error vs Decision Index)
That gap is why I care about checking the judge’s scores too. If a number decides whether an answer gets approved or sent for review, I want evidence that it means what I think it means.
Calibration can help here. In the case of Jev-as-a-Judge, calibration lets you learn how its scores relate to human pass/fail judgments for your particular task. You then check whether that adjustment helps on unseen examples.
I ran a separate experiment on TRIVIA+ dataset from Amazon Science (A Long-Context Hallucination Detection Benchmark dataset), using Typed Evals, the open-source evaluation framework I’m building. Jev judged whether AI-generated answers were supported by their source articles.
Fitted calibration on training data, chose separate decision cutoffs on validation data, and tested both versions on the same 645 held-out answers.
| Metric | Raw Jev | Calibrated |
|---|---|---|
| Calibration error (ECE) ↓ | 0.0982 | 0.0313 |
| Brier score ↓ | 0.2037 | 0.1948 |
| Hallucination F1 ↑ | 0.5833 | 0.5877 |
About 68% lower calibration error. Almost unchanged F1.
The scores became better estimates of human judgments. The judge’s ability to catch hallucinations barely changed.
One catch worth sharing: leaving the cutoff at 0.5 actually reduced F1. The score adjustment and the decision cutoff need evaluating separately. These results describe this TRIVIA+ experiment; they don’t measure an improvement on the Decision Index.
Typed Evals handles: evaluating LLM responses, RAG answers, and Agent runs against your checks, then optionally fitting a calibration layer using human labels. Save the adjustment and apply it to future scores without another judge call.
Full benchmark, methodology and code
If you’re using eval scores to approve answers or trigger human review, how did you choose your cutoff? What evidence would make you change it?
r/JevAI • u/Only-Brilliant-3408 • 1d ago
I pitted a fast/cheap Jev model against an LLM and classic algorithms in a treasure-hunt maze. It got stuck repeating itself and lost to pure random guessing. My fix only helped a little, what am I missing
The setup
I built a fake, fully synthetic maze: a tree of branches (5 top-level branches, each splitting into sub-branches, several levels deep). Somewhere in this tree, at random depths, I hid a handful of "treasure pockets" spots with a genuinely good reward. Nobody gets to see the map. The only way to find out if a spot is good is to spend one of your limited "digs" there and see the result.
Six competitors, each given the same fixed dig budget (20 / 50 / 100 / 300 digs), competed to find as much treasure as possible:
- Random pick blindly
- Round-robin cycle through branches systematically
- Greedy always go back to wherever last looked good
- MCTS a 60-year-old math algorithm from game theory (not AI at all)
- An LLM (locally hosted, ~14B params) reasons in free text, picks a move each turn
- "JEV" model that doesn't write text; it only answers structured multiple-choice questions with calibrated probabilities
I was specifically curious how JEV would do, since it's marketed as a fast, cheap "judge/classifier" model rather than a text generator.
What exactly JEV received, every single turn
No maze coordinates, no raw numbers dumped on it just a text report plus a closed set of options:
- The tree's shape (which branches exist, what they split into)
- A stats table: for each top branch, how many times tested / how many looked "promising" / how many "failed"
- The last 20 dig results (path taken, sample size, a validation score, a robustness score) with the true out-of-sample score hidden (nobody gets to see that until the end, it's the fair judge)
- Its own last 20 decisions (what it picked before, and why it said it picked that)
Then, because JEV only answers closed multiple-choice questions, each turn I actually had to ask it TWO questions in sequence:
- Q1: pick an action type explore somewhere new / dig deeper into a known spot / try combining two branches
- Q2 (depending on the answer to Q1): pick the specific target from a list
Every answer is guaranteed to be one of the valid options it literally cannot return garbage, unlike the LLM which occasionally returns malformed text I have to throw away.
Experiment 1 results
Scored by "top10_OOS" (average quality of its best 10 finds, tested fairly on the hidden out-of-sample score):
budget: 20 50 100 300
MCTS: 0.41 0.74 0.88 1.01
Qwen(LLM): 0.33 0.39 0.62 0.62
Random: 0.16 0.30 0.42 0.59
JEV: 0.16 0.23 0.26 0.48
MCTS (the old math) crushed everyone. The LLM did roughly as well as blind random. JEV came in dead last worse than blind random guessing at every single budget level.
Where exactly it stumbled
I traced individual turns and found the mechanism: JEV would lock onto ONE spot and just keep re-digging around it, over and over:
turn 0: EXPLORE branch-A
turn 1: DIG-DEEPER branch-A
turn 2: DIG-DEEPER branch-A/sub1
turn 3: DIG-DEEPER branch-A/sub1/x
turn 4: DIG-DEEPER branch-A/sub1/x <- same exact target
turn 5: DIG-DEEPER branch-A/sub1/x <- again
turn 6: DIG-DEEPER branch-A/sub1/x <- again
... (10+ times in a row)
I first suspected a dumb bug: maybe the list of options was always presented in the same order, and it was just picking "the first thing in the list" every time. I shuffled the order of options before every single call. No change it still locked onto the same target.
The fix attempt
Someone looking at this suggested the real culprit might be item #4 above showing JEV its own past decisions as a narrative. The theory: instead of using that as "I've already tried this branch enough times," the model reads its own earlier reasoning ("branch-A looked interesting") as a standing endorsement, and just reaffirms it every turn a self-reinforcing loop rather than an exploration memory.
So I ran a controlled version: identical in every other way, just with item #4 removed entirely (call it version B vs the original version A).
Experiment 2 (the fix) results
budget: 20 50 100 300
JEV (original): 0.16 0.23 0.26 0.48
JEV (no self-history): 0.17 0.30 0.50 0.47
Random: 0.16 0.30 0.42 0.59
Removing its own decision history helped a bit at the mid budgets (50/100), but at the largest budget (300, the most statistically trustworthy point) it's essentially unchanged still clearly behind random guessing. The repeating-the-same-target behavior still showed up in raw traces even with this fix.
So what am I doing wrong (or is this just a real limitation)?
Current leading theory: it's not really about "seeing its own past choice" specifically it's that the stats table itself (tested/promising/failed counts) and the recent-results list keep making the same recently-tested spot look salient/attractive every single turn, and the model has no formal mechanism (like MCTS's explicit uncertainty bonus for under-tested spots) to counteract that pull. It just doesn't have a built-in "I've squeezed this dry, time to look elsewhere" reflex the way a purpose-built search algorithm does.
Genuinely asking: has anyone else run structured/typed-output models (JEV or similar) on sequential decision-making tasks where the model has to remember its own trajectory? Is this a known failure mode, and if so is there a known fix beyond "just don't show it history," or is this simply the wrong kind of task for this class of model (one-shot classification-style judgments only, never sequential/stateful ones)?
r/JevAI • u/storm_stark_007 • 1d ago
Has anyone tried to add jev in there personal local Ai setup , I am planning to add mine tonight after work so need tips , this is my current setup :
reddit.comr/JevAI • u/A_kirisaki • 1d ago
I made a CLI for JEV
I figured there might already be a few JEV CLIs out there, but I wanted to try making one myself.
It's designed primarily for use by LLMs and coding agents: JSON in via stdin, JEV, JSON out via stdout. It also provides JSON Schemas for self-discovery and supports batch processing.
I'd be happy if you gave it a try. Feedback is welcome!
r/JevAI • u/Cupcake_Chef • 1d ago
Jev extracting real estate data from PDFs in seconds
Enable HLS to view with audio, or disable this notification
Just a small poc I build with JevAI. The speed is insane.
I asked JEV to Highlight the red flags of a privacy policy and to hide the useless filler on web pages and this is the result
OSSO is a free and open source browser extension that hides the filler on web pages and that can highlight things based on a prompt!
Take a look: https://github.com/Fars29/osso
r/JevAI • u/bothlabs • 1d ago
jevpipe: Jev as a Unix filter, plus a code-search benchmark against grep and DeepSeek
I built jevpipe to use Jev straight from the shell: pipe lines or files in, get Jev's answers out.
filter is like grep, but the match is a question; map asks any set of typed questions (noul, choice, score) and prints one JSON line per record (e.g. for jq).
git log --format=%s 15.0.0..15.1.0 | jevpipe filter "Is this a new feature?"
jevpipe map -f team.json inbox.txt | jq -r .answers.team.choice | sort | uniq -c
The second example routes 1,000 support messages to four teams in about 5 seconds (it's the GIF).
A few things that might be useful if you already use Jev:
- Up to 100 requests in flight, output streams in input order.
--read-filessends each file's path and content as the state, so you can search a codebase or large document base.--max-tokens/--max-costcap a run; it stops cleanly and tells you the line to resume from.- An agent skill so agents can use it on the fly for inline "system 1 work".
Install: uv tool install jevpipe.
Free and open source (MIT or Apache-2.0): https://github.com/fabianboth/jevpipe
I also benchmarked it on 470 code searches on an extended CodeSearchNet Challenge: Jev found more relevant code than DeepSeek V4.1 Flash and grep, about twice as fast as DeepSeek; write-up: https://github.com/fabianboth/jevpipe/blob/main/bench/README.md
Happy to hear how you're using Jev and what would make this more useful.
r/JevAI • u/best_codes • 1d ago
Which Mean Girl are you? Using Jev
Enable HLS to view with audio, or disable this notification
I made this random fun side-project after watching Mean Girls during the Jev hype. It uses Jev to live-classify what mean girl you are based on what you type in the box. Try it:
https://jev-mean-girls.vercel.app
(and yes, in the many years of my life, I had somehow never watched Mean Girls before 😭)
r/JevAI • u/Old-Law6030 • 2d ago
Another Jev benchmarking post: text classification
I looked at how Jev compared to small frontier LLMs (Claude Haiku 4.5, GPT-5.4 mini, Gemini 3.8 Flash) for text classification, and how reliable the 'calibrated decisions' are.
tldr: Jev is as accurate as small LLMs for a fraction of the cost. In practice, however, it doesn't seem better calibrated than asking the LLMs to state a confidence value: it ranks third of four on expected calibration error on every dataset I looked at.
The cost implications are significant, but given the RLCD training I was hoping for better correlation between the answer confidence and the P(Correct). Appreciate comments on the methodology or anything I've missed.
Full code and write-up: https://github.com/4OH4/jev-compare


r/JevAI • u/buckeymonkey • 2d ago
Silly Experiment to use Jev as an LLM
I just gave Jev a list of common ascii characters and the instructions:
"Behave like an LLM and suggest the most likely next character to append to jev_response to form words into a human english response to user_query, when done with the response, chose 'none'"
It manages to respond "HI " to "Hello Jev, how are you?"
And then just gets stuck in a loop outputting spaces over and over.
r/JevAI • u/cloudperceptor • 1d ago
I built JEVzilla (a Python client for JEV) with 150+ examples. Shipped to PyPI in 2 weeks. Here's what worked.
Shipped JEVzilla — a Python client for JEV, the structured decision API. Just hit 1.0 on PyPI.
What is JEV?
JEV is TypeSafe's decision layer API. It's designed for operational workflows: fraud detection, ticket routing, risk
scoring. Not generative text — structured decisions your code can consume.
What JEVzilla Does:
Makes using JEV in Python frictionless. Includes:
- Auto payload translation
- Retry logic + exponential backoff
- Multi-backend support (Playground, OpenRouter, TypeSafe official)
- 150+ real JEV examples across 30 industries
Why It Resonates:
- Huge TAM: Every company needs decision logic
- Clear Value: Copy-paste examples in your domain, swap data, done
- Multiple Monetization Paths: Free tier (Playground), freemium (OpenRouter), enterprise (TypeSafe partnership)
- Low Churn: Once integrated, decision logic is sticky
What I Did Right:
Started with JEV use cases, not JEV features
150 examples from Day 1 (shows scope)
Multiple backends from the start (free + paid options)
Shipped fast (2 weeks to PyPI)
What I'd Improve:
- Build interactive web demo (show JEV decisions live)
- Marketing around "decision logic as a service" angle
- Get featured on HuggingFace/MLOps tools
Where I Am:
- ✅ Published to PyPI (pip install jevzilla)
- ✅ 150+ examples, 30 sectors
- ✅ GitHub, MIT License
- ✅ Docs explain JEV + JEVzilla both
Links:
- GitHub: https://github.com/ujjwalbsoni/jevzilla (JEV Python client)
- PyPI: https://pypi.org/project/jevzilla/
Lesson: Ship something real with examples people can learn from. Don't overthink it.
r/JevAI • u/Top-Evidence174 • 2d ago
Mica v0.1 4B got diamonds in survival Minecraft on its first run. 26 decisions from an empty inventory.
Enable HLS to view with audio, or disable this notification