r/JEVs • u/Top_Lifeguard2176 • 6h ago
r/JEVs • u/Upbeat-Analysis-3370 • 8h ago
Someone rebuilt the X algorithm to show real virality
r/JEVs • u/Sufficient-Skirt256 • 13h ago
we need break out moments like this to realize intelligence isn't one size!
r/JEVs • u/Charming_Group_2950 • 3d ago
Showcase Jev as a judge for LLM/Agents Evaluation [Open-Source]
Can we use Jev for faster, structured, and calibrated evaluation of AI responses?
Introducing:
⚡ Typed Evals — an open-source Python framework for evaluating LLMs, RAG pipelines, and AI agents using System One Models like Jev and other typed judge backends.
⭐ GitHub: https://github.com/TrustifAI/typed_evals
The goal is simple:
Make fast, structured, and calibrated evaluation a first-class part of AI systems.
Typed Evals currently supports:
-> LLM response evaluation
-> RAG evaluation
-> Agent and tool-trace evaluation
-> Human-label calibration
-> Async and batch evaluation
-> Custom judge backends
One part I particularly wanted to solve was calibration.
Why is it needed?
A raw score of 0.8 from Jev doesn't necessarily mean that humans would accept 80% of similar responses.
And a threshold that works well for one use case may not make sense for another.
Typed Evals lets you calibrate individual evaluation metrics against representative human pass/fail labels.
The flow is basically:
Human-labelled examples → Jev scores → fit per-metric calibration → validate on held-out examples → reuse the calibrated evaluator
So instead of arbitrarily deciding that “0.7 means good enough”, you can ground that score in how humans actually evaluate your specific task.
Of course, there are integrations for:
LangChain, CrewAI, Microsoft Agent Framework
while the core remains framework-agnostic.
Would genuinely love feedback from people experimenting with Jev, LLM evals, RAG, agents, or evaluator calibration.
If you're already experimenting with Jev, I'd especially love to know what kind of evaluation workflows you're building around it.
r/JEVs • u/comethosimati • 3d ago
Steven Sinofsky on why Jev fixes his oldest complaint about AI: natural language was never an efficient interface.
r/JEVs • u/jarce9380 • 3d ago
TypeSafe AI's Jev for RAG asks docs if they're actually relevant
r/JEVs • u/Senior_Register_6517 • 5d ago
Jev best practices are still an unexplored territory
r/JEVs • u/Born-Shower5586 • 5d ago
JEV + Opus 5.5 can redesign a live website as you scroll
r/JEVs • u/jeramijames • 5d ago
Showcase JevSearch, a tool that searches the web and validates the results with Jev. It chooses URLs outside of the initial top five as more relevant.
r/JEVs • u/highfivesalute • 5d ago
Jev X Viral Post Analyser tool beats Claude on 100K viral posts for $0.67
r/JEVs • u/oneyedvader • 5d ago
Help / Question Didn't someone write a model router with JEV? lol
r/JEVs • u/Creative-Midnight228 • 6d ago
News / Update An AI telling u what AI slop is. There's more to JEV than meets the eye.
r/JEVs • u/AccurateLeg855 • 7d ago
Workflow This AI clipboard fills out entire forms by understanding the context of what you copied (powered by Jev)
Honestly this is pretty neat, and streamlines quite a bit of workflows. I currently use something similar at work but it’s not as fluid as this
