DEV Community

Cover image for Four hundred distinct strings is not four hundred authors
Adab ul Qayyum
Adab ul Qayyum

Posted on

Four hundred distinct strings is not four hundred authors

The export has a column called Author. A couple of thousand rows, maybe four hundred distinct strings — except nobody knows whether it's four hundred, because "Margaret Atwood" and "margaret atwood" are both in there, and so, somewhere, is "Atwood, Margaret".

In the old system none of this mattered. The column was text, it rendered as text, and nothing ever asked how many authors the catalogue contained.

Now you're moving it somewhere that has entities, and the question has to be answered.

The capacity question is already closed

Worth getting out of the way, because plenty of migration plans spend their time here and it's the wrong place.

A Shopify store gets 128 metaobject definitions on Basic, Shopify and Advanced, 256 on Plus and Enterprise. Each installed app gets its own 128 on top. Every definition holds up to a million entries — raised from the old 64,000 for non-Plus and 128,000 for Plus, with the plan distinction removed entirely. Shopify's own standard definitions don't count against any of it.

So you can have an Author entity. Four hundred of them, or four hundred thousand. You can reference them from products and edit a biography once rather than on every title.

None of that is the hard part.

The hard part is counting

Promoting a column to an entity requires you to know how many entities are in it, and a flat text column doesn't tell you.

Four hundred distinct strings isn't four hundred authors. It's four hundred strings. Some are the same person typed differently. Some are the same person with a middle initial on half their titles. Some are genuinely two different people who share a name.

Counting distinct values gives you a number that looks authoritative and isn't — and the moment you create one metaobject per distinct string, that wrong number becomes the structure of the catalogue.

This is an identity problem with no identifier to work from. The source system never issued one, because it never thought of an author as a thing.

A rule for what to promote

What you can automate is the triage. Which columns are obviously per-product, which are obviously shared, and which can't be judged until a person looks.

A column where nearly every value is unique is a metafield — it's an attribute of the product, and giving each value its own entity creates a thousand records referenced once each
A column where values repeat substantially is a metaobject candidate
A column whose distinct count moves depending on how you normalise it is neither, yet — the number the decision depends on isn't knowable

That third case is the one worth building for, and the useful behaviour is refusal.
`/** A metaobject earns its place when the average value is used more than once.

  • Below this, you have a per-product attribute wearing an entity's clothes. */ const MIN_REUSE = 2;

type Verdict =
| { kind: "skip"; reason: string }
| { kind: "metafield"; distinct: number; reuse: number }
| { kind: "metaobject"; entries: number; reuse: number }
| { kind: "needs-review"; collisions: { normalised: string; spellings: string[] }[] };

const norm = (s: string) => s.trim().toLowerCase().replace(/\s+/g, " ");

/**

  • Decides what a denormalised source column should become in Shopify. *
  • Note what is NOT checked: the per-definition entry ceiling. A metaobject
  • definition holds a million entries, so no real catalogue reaches it, and a
  • check that never fires would imply the ceiling is the risk. It is not. */ export function shouldPromote(values: readonly string[]): Verdict { const present = values.map((v) => v.trim()).filter((v) => v !== ""); if (present.length === 0) return { kind: "skip", reason: "column is empty" };

const groups = new Map>();
for (const v of present) {
const key = norm(v);
const spellings = groups.get(key) ?? new Set();
spellings.add(v);
groups.set(key, spellings);
}

// Spelling collisions come first. While they exist the distinct count is a
// guess, so every number downstream of it would be a guess too.
const collisions = [...groups.entries()]
.filter(([, spellings]) => spellings.size > 1)
.map(([normalised, spellings]) => ({ normalised, spellings: [...spellings].sort() }))
.sort((a, b) => a.normalised.localeCompare(b.normalised));
if (collisions.length > 0) return { kind: "needs-review", collisions };

const distinct = groups.size;
const reuse = Number((present.length / distinct).toFixed(2));
return reuse >= MIN_REUSE
? { kind: "metaobject", entries: distinct, reuse }
: { kind: "metafield", distinct, reuse };
}`

The ordering is the argument: collisions are checked before reuse, because while two spellings of one name are still in the column the distinct count is a guess, and the reuse ratio computed from it would be a guess with a decimal point on it. The function refuses to return a number it cannot stand behind. What it does not check is the entry ceiling, and that omission is deliberate — a definition holds a million entries, so a limit check would never fire and would quietly imply the ceiling is what you should worry about. The test worth reading is the last one: "Stephen King" and "King, Stephen" are reported as two separate entities. That is wrong, the function cannot tell, and making it guess would be worse than leaving it visibly wrong.

Safe normalisation vs. guessing

The line between them is sharper than it looks.

Trimming whitespace is safe. A name with a trailing space and the same name without one are the same string with no information between them. Nothing is lost by collapsing them on the way in.

Lowercasing for comparison is already a judgement. You can match on it, but you can't tell which casing is the one to keep — and picking one silently means the catalogue displays somebody's name the way the import happened to see it first.

Past that it stops being automatable at all. A reordered name. An initial on some titles and not others. A translator credited as a narrator on one record. An imprint that changed its name halfway through the catalogue. All of these are the same entity to a human and different strings to any rule you can write.

The honest architecture surfaces them for review rather than resolving them, and accepts that someone who knows the catalogue has to spend an afternoon on it.

That afternoon is the actual cost of the decision, and it's why the decision gets deferred.

Deciding late costs more than deciding wrong

Which is the trap, because deferring is the expensive option.

Import the column as plain text and everything works. Product pages render, the catalogue is live, and the structural question is still open — but it's now open across the live catalogue rather than across a spreadsheet.

Promoting a text field to an entity afterwards means creating the entities, resolving the duplicates you avoided resolving the first time, rewriting every product's reference, and doing it on data that customers and staff have since edited.

Mapping a column to the wrong structure is recoverable. Mapping it before you've asked what's in it means the recovery happens later, with more rows and an audience.

When this needs an engineer

It doesn't, when the data is genuinely per-product. Most columns in most exports are attributes, not entities, and a metafield is the correct and cheap answer. The guides telling you to use a metaobject for your size chart and a metafield for your SKU are right — for data you're about to create, the reuse question is easy, because you control the answer.

It needs engineering when the data already exists in a shape somebody else chose, at a volume nobody can read. Tens of thousands of products, a column that's probably an entity, and no identifier anywhere in the source to resolve it by. That's a data-modelling job with an audit in front of it, and the audit has to happen before the import rather than after.

Takeaways
Capacity is not the constraint. 128 metaobject definitions, 256 on Plus, and a million entries each — the old 64,000 and 128,000 caps are gone.
A distinct-value count is not an entity count. Four hundred strings may be three hundred and forty people, and creating one metaobject per string makes the wrong number permanent.
Trimming whitespace is safe; choosing a capitalisation is a guess. Normalise the first silently, escalate the second.
Build the triage to refuse. A tool that reports a confident number it cannot stand behind is worse than one that says a person needs to look.
Decide the model before the transformation. Promoting a text column to an entity after go-live means resolving the same duplicates on live data that people have since edited.

Top comments (0)