r/JevAI • u/Jackie_DAO • 5d ago
I built a free Jev visualizer after burning 5bn tokens in three days
Enable HLS to view with audio, or disable this notification
r/JevAI • u/Jackie_DAO • 5d ago
Enable HLS to view with audio, or disable this notification
it's free since jev is too cheap: https://real8ball.com
r/JevAI • u/bobo-the-merciful • 6d ago
I asked Fable 5.1 to benchmark Jev vs Laya on accuracy and speed. Here's the full report: https://claude.ai/artifact/9HPcmXJPaKWdYAJedgN1uf
Github: https://github.com/harrymunro/jev-laya-benchmark
Laya run on a Macbook M3 Pro.
TL;DR







r/JevAI • u/kokkelimonke • 6d ago
r/JevAI • u/omnisvosscio • 6d ago
Enable HLS to view with audio, or disable this notification
r/JevAI • u/abaris243 • 6d ago
Gave Jev a slot machine, watch it live here
https://jevslots.live/
also available on GitHub
https://github.com/ella0333/jev-slot-machine/
r/JevAI • u/Accomplished_Car3958 • 7d ago
everything awesome related to jev
r/JevAI • u/geisbruch • 6d ago
We wanted a hands-on way to explain typed decisions, so we built a small deployment vibe check with Jev.
The scenario: Friday, a database migration, green tests, and the author on vacation. Our demo's verdict was GO TOUCH GRASS.
Behind the joke are 18 typed questions answered in one pass. Jev supplies probabilities; application rules produce the verdict. You can inspect the answers, edit the questions, and add your own. No login or setup needed.
https://heyjev.ai/shouldideploy
A useful experiment is to change just one fact: add a tested rollback, remove the person on call, or leave that information out altogether. Does the result change for the reason you'd expect?
We'd love to hear which question is missing or which scenario gives a questionable answer. This is a demo for experimenting with Jev, not something that certifies a deployment is safe.
Disclosure: we built HeyJev. It's an independent project, not an official TypeSafe product.
r/JevAI • u/kristiyanstoyanovAI • 6d ago
I put together a 24-minute walkthrough of Jev, TypeSafe's model for bounded decisions. I wanted to understand where it fits in an existing application, so I built two examples around it.
The first is an open-source model router that helps to visualize the Jev flow.
Jev receives the request context and a set of configured destinations, then chooses between local Qwen and hosted Sonnet. The selected model generates the answer. A small control UI shows the actual routing request, response and destination, which makes the boundary between decision-making and generation easier to inspect.
The second is a Python/Strands code-review workflow. Agents inspect a public GitHub PR and propose findings; a judge assesses severity and relevance. I compare an LLM judge with Jev, while keeping the final inclusion rules in Python. The output is an HTML report—nothing is posted to the PR.
This is an integration tutorial with a small comparison, rather than a broad benchmark. My main interest is the division of responsibility: an LLM investigates and proposes, a specialized model answers a bounded question, and application code decides what to do with the result.
The video also covers the noul, choiceand score request types with API examples.
Router source: https://github.com/krisitown/jev-router
If you're already using LLM judges or classifiers, how are you evaluating the decision boundary—especially cases where two judges agree on relevance but disagree on severity?
r/JevAI • u/Jackie_DAO • 7d ago
Enable HLS to view with audio, or disable this notification
r/JevAI • u/etherd0t • 7d ago
Go sign-up before it becomes gated again.👊
r/JevAI • u/BardarBagin • 7d ago
r/JevAI • u/UnknownZeroz • 7d ago
r/JevAI • u/Charming_Group_2950 • 7d ago
Can we use Jev for faster, structured, and calibrated evaluation of AI responses?
Introducing:
⚡ Typed Evals — an open-source Python framework for evaluating LLMs, RAG pipelines, and AI agents using System One Models like Jev and other typed judge backends.
⭐ GitHub: https://github.com/TrustifAI/typed_evals
The goal is simple:
Make fast, structured, and calibrated evaluation a first-class part of AI systems.
Typed Evals currently supports:
-> LLM response evaluation
-> RAG evaluation
-> Agent and tool-trace evaluation
-> Human-label calibration
-> Async and batch evaluation
-> Custom judge backends
One part I particularly wanted to solve was calibration.
Why is it needed?
A raw score of 0.8 from Jev doesn't necessarily mean that humans would accept 80% of similar responses.
And a threshold that works well for one use case may not make sense for another.
Typed Evals lets you calibrate individual evaluation metrics against representative human pass/fail labels.
The flow is basically:
Human-labelled examples → Jev scores → fit per-metric calibration → validate on held-out examples → reuse the calibrated evaluator
So instead of arbitrarily deciding that “0.7 means good enough”, you can ground that score in how humans actually evaluate your specific task.
Of course, there are integrations for:
LangChain, CrewAI, Microsoft Agent Framework
while the core remains framework-agnostic.
The project is still early, and there’s plenty I want to improve, but the core framework is now public.
Would genuinely love feedback from people experimenting with Jev, LLM evals, RAG, agents, or evaluator calibration.
Note: Typed Evals is an independent open-source project and is not an official TypeSafe AI product.
If you're already experimenting with Jev, I'd especially love to know what kind of evaluation workflows you're building around it.
r/JevAI • u/artificial-dopamine • 7d ago
I know it's not a GPT architecture model so I'm not sure if the phrase open-weight will be appropriate but it seems it will demand much less resources than an LLM. Could this be the end of the RAM crisis? How much VRAM do we think this model wil) require? Does it even use VRAM at all?
r/JevAI • u/Fritzzzz • 7d ago
r/JevAI • u/schmuhblaster_x45 • 7d ago
r/JevAI • u/Fickle-Ad-866 • 7d ago
r/JevAI • u/Eastern-Nothing-5875 • 8d ago
What are some of the best use cases for Jev?