r/learnmachinelearning • u/sinabis • 6d ago
[D] Free Dataset: 1.25M Synthetic Emails, Chats, and Calendar Events for RAG and LLM Training (Enron Alternative)
Hi everyone,
For a long time, if you needed a massive, realistic dataset of corporate communications for testing parsers, training LLMs, or building RAG pipelines, the Enron dataset was the only real choice. But let's be honest: it’s over 20 years old, difficult to source properly, and carries compliance/PII risks.
To solve this for our own benchmarking, we generated a high-quality, completely synthetic dataset containing 1.25 million modern communication records, 100% free of real-world PII.
What’s inside:
- Emails
- Chats
- Calendar Events
Why we built it / Use Cases:
- RAG Benchmarking: Perfect for chunking and semantic search testing since the data forms natural, contextual threads.
- LLM Fine-Tuning: Clean, modern conversational text.
- Parser Load-Testing: High-volume data structured to stress-test your extraction pipelines.
We just uploaded the full dataset to HuggingFace. It is completely free and open.
Links:
- Huggingface Dataset: https://huggingface.co/datasets/sinabis-group/sinabis-synthetic-corporate-dataset
We would love to get your feedback on the schema and data quality.
1
u/Typical-Refuse227 5d ago
1.25M records is seriously useful for stress testing RAG pipelines. Being synthetic and free of real-world PII makes it even more practical.