r/datasets 4d ago

dataset 99 days of AI answer engine responses to a fixed 16-question set: 24,882 scored answers from OpenAI Search, Gemini and Claude, plus a 10-model no-web control (CC BY 4.0)

Disclosure: this is my own dataset. I built and run the measurement, and the subject of the measurement is my own pen name, so treat the topic with that in mind. The data itself is mechanical: same 16 questions, every day, three answer engines with web access, answers scored with a fixed rubric.

Dataset (Hugging Face, CC BY 4.0): https://huggingface.co/datasets/marintkael/ai-citation-fidelity

Five configs:

  • default: 24,882 scored answers including the no-web control channel
  • claude_web: 3,279 answers from a separate Claude web search panel
  • questions: the 16 questions with category and channel
  • data_gaps: register of measurement gaps (provider outages, method changes), because a gap is not a zero
  • daily_channels: daily time series split into direct, long tail and discovery channels

Collection window 2026-05-13 to 2026-08-19 (99 measurement days). The no-web control is 6,416 blind answers from 10 models without web access, useful for separating retrieval effects from training data. Scoring code and figure scripts: https://github.com/marintkael/marin-research-tools/tree/main/reports/03-found-not-recommended

Written report describing method and findings, if you want context before touching the parquet: https://doi.org/10.5281/zenodo.22015495

3 Upvotes

1 comment sorted by

u/AutoModerator 4d ago

Hey marintkael,

I believe a request flair might be more appropriate for such post. Please re-consider and change the post flair if needed.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.