DEV Community

Cover image for How to Generate Realistic Test Data Without Using Real Customer Data
Genory Team for Genory

Posted on

How to Generate Realistic Test Data Without Using Real Customer Data

Realistic test data is one of those things that seems simple until you actually need it.

A signup form might only require a name and an email address. But a real application often needs much more:

  • names
  • addresses
  • phone numbers
  • dates
  • UUIDs
  • company information
  • account-format data
  • custom fields
  • relationships between fields

The obvious shortcut is to copy a few rows from production.

That is also one of the worst habits a development team can build.

In this article, we'll look at a safer and more useful approach: generating synthetic test data designed specifically for development and QA.

Why production data should stay out of your test environment

Using real customer information may make test data look realistic, but it introduces unnecessary risk.

Development, staging and demo environments often have different security controls than production. Data may end up in:

  • local databases
  • screenshots
  • bug reports
  • log files
  • test exports
  • developer laptops
  • temporary environments

A much better principle is:

Test the structure and behavior of your application without copying the people behind the data.

Synthetic test data gives you values that resemble the data your application expects while remaining independent from actual customer records.

Random data is not always good test data

Generating random strings is easy.

Generating useful test data is harder.

Consider an address form.

This is technically random:

Name: Xkqpd Azzw
City: 48291
Phone: foo-bar
Enter fullscreen mode Exit fullscreen mode

But it doesn't help much when testing a real user interface.

A more useful fixture might look like:

Name: Anna Schneider
Email: anna.schneider@example.com
Country: DE
City: Hamburg
Postal code: 20095
Enter fullscreen mode Exit fullscreen mode

The goal is not to create a real person.

The goal is to produce data that behaves like the type of input your application is designed to process.

Keep related fields related

One common mistake is generating every column independently.

Imagine this row:

{
  "country": "DE",
  "postalCode": "SW1A 1AA",
  "phone": "+81...",
  "city": "Toronto"
}
Enter fullscreen mode Exit fullscreen mode

Every individual value may look plausible, but the record as a whole is useless for many tests.

For profile-style fixtures, related fields should share context.

Names, addresses, phone formats and country settings should make sense together whenever your test actually depends on that relationship.

For tests where relationships don't matter, you can deliberately generate fields independently.

The important part is making that choice intentionally.

Use reserved domains for test email addresses

Email addresses are another surprisingly easy source of trouble.

Avoid generating random addresses on real domains such as:

randomperson@gmail.com
Enter fullscreen mode Exit fullscreen mode

You don't know whether that address actually belongs to someone.

For test fixtures, domains such as example.com are much safer:

maria.schmidt@example.com
Enter fullscreen mode Exit fullscreen mode

They clearly communicate that the address is test data and avoid accidentally involving real users.

Make test datasets reproducible

Random data is useful for exploration.

Repeatable data is useful for debugging.

Suppose a test fails only when a particular dataset is generated. If every execution creates completely different values, reproducing the problem becomes harder.

A seeded generator solves this.

Conceptually:

seed = checkout-regression-42
Enter fullscreen mode Exit fullscreen mode

Using the same generator configuration and seed can reproduce the same dataset.

This is especially useful for:

  • regression tests
  • imports
  • API fixtures
  • UI snapshots
  • bug reproduction
  • CI pipelines

Use random datasets when you want variation.

Use seeded datasets when you want repeatability.

Test boundaries, not just happy paths

Synthetic data becomes much more valuable when you stop treating it as filler.

Instead, design datasets around scenarios.

For example:

Scenario 1: normal signup
Scenario 2: very long name
Scenario 3: missing optional address field
Scenario 4: minimum allowed number
Scenario 5: maximum allowed number
Scenario 6: leap-day date
Scenario 7: duplicate identifier
Enter fullscreen mode Exit fullscreen mode

The best test dataset is rarely the largest one.

It is the dataset that deliberately exercises the assumptions in your application.

Structured identifiers need special handling

Some values have rules beyond simple formatting.

Examples include:

  • UUIDs
  • IBANs
  • card-number formats
  • IMEI numbers
  • MAC addresses

For these, a random string with the right length is often not enough.

A UUID should follow the appropriate UUID format.

An IBAN used to test a validator may need the correct country structure and checksum.

A card-number fixture may need to satisfy the Luhn algorithm.

That still does not mean the generated value represents a real account, device or payment method.

Format validity and real-world existence are two completely different things.

Generate only the fields you need

Another mistake is generating huge fake profiles for every test.

If you're testing an import with these columns:

customer_name
email
order_total
status
Enter fullscreen mode Exit fullscreen mode

you probably don't need:

phone
street
company
job_title
iban
username
date_of_birth
Enter fullscreen mode Exit fullscreen mode

Smaller datasets are easier to understand and debug.

A schema-first approach works well:

[
  { "name": "customer", "type": "fullName" },
  { "name": "email", "type": "email" },
  { "name": "total", "type": "decimal" },
  { "name": "status", "type": "choice" }
]
Enter fullscreen mode Exit fullscreen mode

Generate the minimum dataset that exercises the behavior you want to test.

Automate test-data generation when it makes sense

Manual generators are useful while developing and exploring.

Once a test-data workflow becomes repetitive, an API can be more practical.

For example, a profile request could look like this:

curl --request POST "https://genory.dev/api/profile" \
  --header "Authorization: Bearer $GENORY_API_KEY" \
  --header "Content-Type: application/json" \
  --data '{
    "country": "DE",
    "amount": 1,
    "fields": ["firstName", "lastName", "email"]
  }'
Enter fullscreen mode Exit fullscreen mode

Now a fixture can be generated from a script, CI job or development tool instead of being copied manually.

Whatever service you use, keep API keys in environment variables or a secret manager rather than committing them to your repository.

A practical workflow

A simple process works surprisingly well:

  1. Define the behavior you want to test.
  2. Decide which fields actually matter.
  3. Choose the country or format constraints.
  4. Add edge cases intentionally.
  5. Use synthetic rather than production data.
  6. Use a seed when reproducibility matters.
  7. Generate a small dataset first.
  8. Scale only when the small test works.

For quick experiments, you can build synthetic profiles and custom datasets with Genory.

Genory includes a Test Data Generator, schema-based dataset tools, UUID utilities and format-specific generators.

If you're automating the process, the current API reference is available in the Genory developer documentation.

Final thought

Good test data isn't data that looks impressive.

It's data that exposes assumptions.

Synthetic data gives developers a way to build realistic fixtures without turning production customer information into development material.

Generate less data, make it intentional, and design every dataset around the behavior you're actually trying to test.

Top comments (1)

Collapse
 
elijahbrown profile image
Elijah Brown •

Your reserved-domain rule for email has a phone version that's easy to miss, because a random number in the right shape can belong to a real person. Ofcom sets aside London 020 7946 0xxx for TV and radio drama and NANPA reserves 555-0100 to 555-0199 as fictitious, so +44 20 7946 0958 or +1 212 555 0123 make safe UK and US fixtures. They're also well formed (libphonenumber treats the London one as a valid fixed-line number), so they go down the same code path as a real number, which makes them handy for checking what your own signup form decides to do with them.