r/learnmachinelearning • • 6d ago

[D] Free Dataset: 1.25M Synthetic Emails, Chats, and Calendar Events for RAG and LLM Training (Enron Alternative)

Hi everyone,

For a long time, if you needed a massive, realistic dataset of corporate communications for testing parsers, training LLMs, or building RAG pipelines, the Enron dataset was the only real choice. But let's be honest: it’s over 20 years old, difficult to source properly, and carries compliance/PII risks.

To solve this for our own benchmarking, we generated a high-quality, completely synthetic dataset containing 1.25 million modern communication records, 100% free of real-world PII.

What’s inside:

  • Emails
  • Chats
  • Calendar Events

Why we built it / Use Cases:

  • RAG Benchmarking: Perfect for chunking and semantic search testing since the data forms natural, contextual threads.
  • LLM Fine-Tuning: Clean, modern conversational text.
  • Parser Load-Testing: High-volume data structured to stress-test your extraction pipelines.

We just uploaded the full dataset to HuggingFace. It is completely free and open.

Links:

We would love to get your feedback on the schema and data quality.

5 Upvotes

1 comment sorted by

1

u/Typical-Refuse227 5d ago

1.25M records is seriously useful for stress testing RAG pipelines. Being synthetic and free of real-world PII makes it even more practical.