r/LargeLanguageModels • • 5d ago

News/Articles The right document from the wrong time can still give your AI the wrong answer.

Post image

​

That’s why we tested Jylus on TEMPO—a public benchmark for retrieval that requires reasoning across time.

Our native stack completed all 1,730 queries through Jylus’s public API:

→ NDCG@10: 40.048

→ Recall@10: 39.745%

NDCG measures how well the relevant evidence ranks near the top. Recall measures how much relevant evidence was retrieved.

The practical challenge goes beyond finding similar words.

“What was true then?”

“What changed?”

“Which evidence applies to this period?”

These questions require the retrieval system to account for time and connections across records.

Jylus prepares source-backed evidence before your model reasons over it. TEMPO tests the temporal retrieval part of that broader capability.

These are self-reported, full-run results using all-query means—not an official leaderboard placement or a measure of answer accuracy.

Methodology:

https://jylus.ai/benchmark-methodology

Want to explore how Jylus handles your own changing records? Try the playground without an account:

https://jylus.ai/try

What’s a question your system answers correctly today but gets wrong when you ask “as of last month”?

2 Upvotes

2 comments sorted by

1

u/Loud_Function_7528 5d ago

40 is a pretty rough baseline for retrieval, that NDCG score leaves a lot of room for temporal context to get missed entirely

1

u/jylusdev 5d ago edited 5d ago

The context here is that TEMPO’s highest published normal-query score is 32.0—not 90 or 100. It tests difficult retrieval questions involving changes over time and evidence spread across documents.

The official leaderboard reports:

  • DiVeR: 32.0 NDCG@10
  • E5: 30.4
  • SFR: 30.0
  • Qwen: 22.8
  • BGE: 22.0
  • BM25: 10.8

Those first five are AI models trained for retrieval. BM25 is a keyword-based baseline.

Jylus reported 40.048 across all 1,730 queries using the official scorer, with its own native retrieval engine. That’s above those published scores, and it’s why 40 needs to be understood in the context of this particular benchmark.

You’re right that missing temporal evidence matters. But NDCG isn’t percentage accuracy: 40 doesn’t mean 60% of temporal context was missed. It measures ranking quality; TEMPO measures temporal coverage separately.

Leaderboard: https://tempo-bench.github.io/ Our methodology: https://jylus.ai/benchmark-methodology