<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Genory</title>
    <description>The latest articles on DEV Community by Genory (genory).</description>
    <link>https://dev.to/genory</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Forganization%2Fprofile_image%2F15087%2F1237dd69-4d8e-4187-9a23-523c23d14fed.png</url>
      <title>DEV Community: Genory</title>
      <link>https://dev.to/genory</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/genory"/>
    <language>en</language>
    <item>
      <title>How to Generate Realistic Test Data Without Using Real Customer Data</title>
      <dc:creator>Genory Team</dc:creator>
      <pubDate>Mon, 05 Oct 2026 09:00:27 +0000</pubDate>
      <link>https://dev.to/genory/how-to-generate-realistic-test-data-without-using-real-customer-data-11l5</link>
      <guid>https://dev.to/genory/how-to-generate-realistic-test-data-without-using-real-customer-data-11l5</guid>
      <description>&lt;p&gt;Realistic test data is one of those things that seems simple until you actually need it.&lt;/p&gt;

&lt;p&gt;A signup form might only require a name and an email address. But a real application often needs much more:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;names&lt;/li&gt;
&lt;li&gt;addresses&lt;/li&gt;
&lt;li&gt;phone numbers&lt;/li&gt;
&lt;li&gt;dates&lt;/li&gt;
&lt;li&gt;UUIDs&lt;/li&gt;
&lt;li&gt;company information&lt;/li&gt;
&lt;li&gt;account-format data&lt;/li&gt;
&lt;li&gt;custom fields&lt;/li&gt;
&lt;li&gt;relationships between fields&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The obvious shortcut is to copy a few rows from production.&lt;/p&gt;

&lt;p&gt;That is also one of the worst habits a development team can build.&lt;/p&gt;

&lt;p&gt;In this article, we'll look at a safer and more useful approach: generating synthetic test data designed specifically for development and QA.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why production data should stay out of your test environment
&lt;/h2&gt;

&lt;p&gt;Using real customer information may make test data look realistic, but it introduces unnecessary risk.&lt;/p&gt;

&lt;p&gt;Development, staging and demo environments often have different security controls than production. Data may end up in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;local databases&lt;/li&gt;
&lt;li&gt;screenshots&lt;/li&gt;
&lt;li&gt;bug reports&lt;/li&gt;
&lt;li&gt;log files&lt;/li&gt;
&lt;li&gt;test exports&lt;/li&gt;
&lt;li&gt;developer laptops&lt;/li&gt;
&lt;li&gt;temporary environments&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A much better principle is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Test the structure and behavior of your application without copying the people behind the data.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Synthetic test data gives you values that resemble the data your application expects while remaining independent from actual customer records.&lt;/p&gt;

&lt;h2&gt;
  
  
  Random data is not always good test data
&lt;/h2&gt;

&lt;p&gt;Generating random strings is easy.&lt;/p&gt;

&lt;p&gt;Generating useful test data is harder.&lt;/p&gt;

&lt;p&gt;Consider an address form.&lt;/p&gt;

&lt;p&gt;This is technically random:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Name: Xkqpd Azzw
City: 48291
Phone: foo-bar
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But it doesn't help much when testing a real user interface.&lt;/p&gt;

&lt;p&gt;A more useful fixture might look like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Name: Anna Schneider
Email: anna.schneider@example.com
Country: DE
City: Hamburg
Postal code: 20095
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is not to create a real person.&lt;/p&gt;

&lt;p&gt;The goal is to produce data that behaves like the type of input your application is designed to process.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep related fields related
&lt;/h2&gt;

&lt;p&gt;One common mistake is generating every column independently.&lt;/p&gt;

&lt;p&gt;Imagine this row:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"country"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"DE"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"postalCode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"SW1A 1AA"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"phone"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"+81..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"city"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Toronto"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every individual value may look plausible, but the record as a whole is useless for many tests.&lt;/p&gt;

&lt;p&gt;For profile-style fixtures, related fields should share context.&lt;/p&gt;

&lt;p&gt;Names, addresses, phone formats and country settings should make sense together whenever your test actually depends on that relationship.&lt;/p&gt;

&lt;p&gt;For tests where relationships don't matter, you can deliberately generate fields independently.&lt;/p&gt;

&lt;p&gt;The important part is making that choice intentionally.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use reserved domains for test email addresses
&lt;/h2&gt;

&lt;p&gt;Email addresses are another surprisingly easy source of trouble.&lt;/p&gt;

&lt;p&gt;Avoid generating random addresses on real domains such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;randomperson@gmail.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You don't know whether that address actually belongs to someone.&lt;/p&gt;

&lt;p&gt;For test fixtures, domains such as &lt;code&gt;example.com&lt;/code&gt; are much safer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;maria.schmidt@example.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;They clearly communicate that the address is test data and avoid accidentally involving real users.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make test datasets reproducible
&lt;/h2&gt;

&lt;p&gt;Random data is useful for exploration.&lt;/p&gt;

&lt;p&gt;Repeatable data is useful for debugging.&lt;/p&gt;

&lt;p&gt;Suppose a test fails only when a particular dataset is generated. If every execution creates completely different values, reproducing the problem becomes harder.&lt;/p&gt;

&lt;p&gt;A seeded generator solves this.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;seed = checkout-regression-42
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Using the same generator configuration and seed can reproduce the same dataset.&lt;/p&gt;

&lt;p&gt;This is especially useful for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;regression tests&lt;/li&gt;
&lt;li&gt;imports&lt;/li&gt;
&lt;li&gt;API fixtures&lt;/li&gt;
&lt;li&gt;UI snapshots&lt;/li&gt;
&lt;li&gt;bug reproduction&lt;/li&gt;
&lt;li&gt;CI pipelines&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use random datasets when you want variation.&lt;/p&gt;

&lt;p&gt;Use seeded datasets when you want repeatability.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test boundaries, not just happy paths
&lt;/h2&gt;

&lt;p&gt;Synthetic data becomes much more valuable when you stop treating it as filler.&lt;/p&gt;

&lt;p&gt;Instead, design datasets around scenarios.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Scenario 1: normal signup
Scenario 2: very long name
Scenario 3: missing optional address field
Scenario 4: minimum allowed number
Scenario 5: maximum allowed number
Scenario 6: leap-day date
Scenario 7: duplicate identifier
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The best test dataset is rarely the largest one.&lt;/p&gt;

&lt;p&gt;It is the dataset that deliberately exercises the assumptions in your application.&lt;/p&gt;

&lt;h2&gt;
  
  
  Structured identifiers need special handling
&lt;/h2&gt;

&lt;p&gt;Some values have rules beyond simple formatting.&lt;/p&gt;

&lt;p&gt;Examples include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;UUIDs&lt;/li&gt;
&lt;li&gt;IBANs&lt;/li&gt;
&lt;li&gt;card-number formats&lt;/li&gt;
&lt;li&gt;IMEI numbers&lt;/li&gt;
&lt;li&gt;MAC addresses&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For these, a random string with the right length is often not enough.&lt;/p&gt;

&lt;p&gt;A UUID should follow the appropriate UUID format.&lt;/p&gt;

&lt;p&gt;An IBAN used to test a validator may need the correct country structure and checksum.&lt;/p&gt;

&lt;p&gt;A card-number fixture may need to satisfy the Luhn algorithm.&lt;/p&gt;

&lt;p&gt;That still does &lt;strong&gt;not&lt;/strong&gt; mean the generated value represents a real account, device or payment method.&lt;/p&gt;

&lt;p&gt;Format validity and real-world existence are two completely different things.&lt;/p&gt;

&lt;h2&gt;
  
  
  Generate only the fields you need
&lt;/h2&gt;

&lt;p&gt;Another mistake is generating huge fake profiles for every test.&lt;/p&gt;

&lt;p&gt;If you're testing an import with these columns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;customer_name
email
order_total
status
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;you probably don't need:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;phone
street
company
job_title
iban
username
date_of_birth
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Smaller datasets are easier to understand and debug.&lt;/p&gt;

&lt;p&gt;A schema-first approach works well:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"customer"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"fullName"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"email"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"email"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"total"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"decimal"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"choice"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Generate the minimum dataset that exercises the behavior you want to test.&lt;/p&gt;

&lt;h2&gt;
  
  
  Automate test-data generation when it makes sense
&lt;/h2&gt;

&lt;p&gt;Manual generators are useful while developing and exploring.&lt;/p&gt;

&lt;p&gt;Once a test-data workflow becomes repetitive, an API can be more practical.&lt;/p&gt;

&lt;p&gt;For example, a profile request could look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;--request&lt;/span&gt; POST &lt;span class="s2"&gt;"https://genory.dev/api/profile"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--header&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$GENORY_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--header&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--data&lt;/span&gt; &lt;span class="s1"&gt;'{
    "country": "DE",
    "amount": 1,
    "fields": ["firstName", "lastName", "email"]
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now a fixture can be generated from a script, CI job or development tool instead of being copied manually.&lt;/p&gt;

&lt;p&gt;Whatever service you use, keep API keys in environment variables or a secret manager rather than committing them to your repository.&lt;/p&gt;

&lt;h2&gt;
  
  
  A practical workflow
&lt;/h2&gt;

&lt;p&gt;A simple process works surprisingly well:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Define the behavior you want to test.&lt;/li&gt;
&lt;li&gt;Decide which fields actually matter.&lt;/li&gt;
&lt;li&gt;Choose the country or format constraints.&lt;/li&gt;
&lt;li&gt;Add edge cases intentionally.&lt;/li&gt;
&lt;li&gt;Use synthetic rather than production data.&lt;/li&gt;
&lt;li&gt;Use a seed when reproducibility matters.&lt;/li&gt;
&lt;li&gt;Generate a small dataset first.&lt;/li&gt;
&lt;li&gt;Scale only when the small test works.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For quick experiments, you can build synthetic profiles and custom datasets with &lt;a href="https://genory.dev/tools/test-data" rel="noopener noreferrer"&gt;Genory&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Genory includes a &lt;a href="https://genory.dev/generate" rel="noopener noreferrer"&gt;Test Data Generator&lt;/a&gt;, schema-based dataset tools, UUID utilities and format-specific generators.&lt;/p&gt;

&lt;p&gt;If you're automating the process, the current API reference is available in the &lt;a href="https://genory.dev/docs" rel="noopener noreferrer"&gt;Genory developer documentation&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final thought
&lt;/h2&gt;

&lt;p&gt;Good test data isn't data that looks impressive.&lt;/p&gt;

&lt;p&gt;It's data that exposes assumptions.&lt;/p&gt;

&lt;p&gt;Synthetic data gives developers a way to build realistic fixtures without turning production customer information into development material.&lt;/p&gt;

&lt;p&gt;Generate less data, make it intentional, and design every dataset around the behavior you're actually trying to test.&lt;/p&gt;

</description>
      <category>testing</category>
      <category>webdev</category>
      <category>api</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
