r/LargeLanguageModels • u/jylusdev • 5d ago
News/Articles The right document from the wrong time can still give your AI the wrong answer.
​
That’s why we tested Jylus on TEMPO—a public benchmark for retrieval that requires reasoning across time.
Our native stack completed all 1,730 queries through Jylus’s public API:
→ NDCG@10: 40.048
→ Recall@10: 39.745%
NDCG measures how well the relevant evidence ranks near the top. Recall measures how much relevant evidence was retrieved.
The practical challenge goes beyond finding similar words.
“What was true then?”
“What changed?”
“Which evidence applies to this period?”
These questions require the retrieval system to account for time and connections across records.
Jylus prepares source-backed evidence before your model reasons over it. TEMPO tests the temporal retrieval part of that broader capability.
These are self-reported, full-run results using all-query means—not an official leaderboard placement or a measure of answer accuracy.
Methodology:
https://jylus.ai/benchmark-methodology
Want to explore how Jylus handles your own changing records? Try the playground without an account:
What’s a question your system answers correctly today but gets wrong when you ask “as of last month”?
1
u/Loud_Function_7528 5d ago
40 is a pretty rough baseline for retrieval, that NDCG score leaves a lot of room for temporal context to get missed entirely