<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sinabis</title>
    <description>The latest articles on DEV Community by Sinabis (@sinabis).</description>
    <link>https://dev.to/sinabis</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4139450%2Fa4b6367c-5dc8-47c2-a5fc-fad502d58157.png</url>
      <title>DEV Community: Sinabis</title>
      <link>https://dev.to/sinabis</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sinabis"/>
    <language>en</language>
    <item>
      <title>Beyond Enron: Why We Built a Free 1.25M Synthetic Email &amp; Chat Dataset for RAG and LLMs</title>
      <dc:creator>Sinabis</dc:creator>
      <pubDate>Wed, 23 Sep 2026 12:46:47 +0000</pubDate>
      <link>https://dev.to/sinabis/beyond-enron-why-we-built-a-free-125m-synthetic-email-chat-dataset-for-rag-and-llms-33h2</link>
      <guid>https://dev.to/sinabis/beyond-enron-why-we-built-a-free-125m-synthetic-email-chat-dataset-for-rag-and-llms-33h2</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnehthnox2wfmqresk2uy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnehthnox2wfmqresk2uy.png" alt=" " width="800" height="354"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you've ever tried to build an email parser, fine-tune an LLM on corporate communications, or benchmark a RAG (Retrieval-Augmented Generation) pipeline, you've likely run into the infamous &lt;strong&gt;Enron dataset&lt;/strong&gt;. &lt;/p&gt;

&lt;p&gt;For over two decades, Enron has been the undisputed "gold standard" for open communication data. But let's be honest with ourselves: &lt;strong&gt;it’s outdated, legally complicated, and filled with massive privacy risks (PII).&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;We faced this exact bottleneck while stress-testing our own data pipelines and AI systems. Real corporate data is locked behind strict NDAs and privacy laws, and Enron just wasn't cutting it anymore for modern tech stacks.&lt;/p&gt;

&lt;p&gt;So, we decided to fix it. Today, we are releasing a completely free, open-source, and high-quality synthetic dataset containing &lt;strong&gt;1.25 million modern communication records&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  📊 What's Inside the Dataset?
&lt;/h2&gt;

&lt;p&gt;Unlike older datasets that only focus on raw email text, we wanted to mirror how modern teams actually communicate. The 1.25M records are split across three major pillars:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Emails&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Chats&lt;/strong&gt; &lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Calendar Events&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Best of all: It is &lt;strong&gt;100% free of real-world PII (Personally Identifiable Information)&lt;/strong&gt;, making it fully compliant and safe to use in any environment.&lt;/p&gt;




&lt;h2&gt;
  
  
  🛠️ Get Started Right Now
&lt;/h2&gt;

&lt;p&gt;The dataset is fully open, hosted on HuggingFace, and ready for you to dive in. The documentation and schema details are already live in our repository's README.&lt;/p&gt;

&lt;p&gt;👉 &lt;strong&gt;&lt;a href="https://huggingface.co/datasets/sinabis-group/sinabis-synthetic-corporate-dataset" rel="noopener noreferrer"&gt;https://huggingface.co/datasets/sinabis-group/sinabis-synthetic-corporate-dataset&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We would love to hear from the community! How are you currently solving the lack of open text datasets for communication?&lt;/p&gt;

&lt;p&gt;Drop your thoughts, feedback, or questions in the comments below!&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>machinelearning</category>
    </item>
  </channel>
</rss>
