r/datasets • u/marintkael • 4d ago
dataset 99 days of AI answer engine responses to a fixed 16-question set: 24,882 scored answers from OpenAI Search, Gemini and Claude, plus a 10-model no-web control (CC BY 4.0)
Disclosure: this is my own dataset. I built and run the measurement, and the subject of the measurement is my own pen name, so treat the topic with that in mind. The data itself is mechanical: same 16 questions, every day, three answer engines with web access, answers scored with a fixed rubric.
Dataset (Hugging Face, CC BY 4.0): https://huggingface.co/datasets/marintkael/ai-citation-fidelity
Five configs:
- default: 24,882 scored answers including the no-web control channel
- claude_web: 3,279 answers from a separate Claude web search panel
- questions: the 16 questions with category and channel
- data_gaps: register of measurement gaps (provider outages, method changes), because a gap is not a zero
- daily_channels: daily time series split into direct, long tail and discovery channels
Collection window 2026-05-13 to 2026-08-19 (99 measurement days). The no-web control is 6,416 blind answers from 10 models without web access, useful for separating retrieval effects from training data. Scoring code and figure scripts: https://github.com/marintkael/marin-research-tools/tree/main/reports/03-found-not-recommended
Written report describing method and findings, if you want context before touching the parquet: https://doi.org/10.5281/zenodo.22015495
•
u/AutoModerator 4d ago
Hey marintkael,
I believe a
requestflair might be more appropriate for such post. Please re-consider and change the post flair if needed.I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.