Shopify apps that offer both checkout-time discounts and storefront widgets face an architectural problem that is rarely discussed systematically: the Function runtime and the Theme Extension runtime live in two isolated data planes that can't directly talk to each other. Your discount rules exist in one world, your storefront widgets live in the other, and keeping them consistent is a surprisingly hard engineering problem.
We built Markdo with 18 Liquid template widgets that display discount information on storefronts — tier tables, flash sale countdowns, cart goals, spin-to-win wheels, and more. This post covers the data sync architecture, the i18n bridge pattern for Liquid-to-JavaScript translation, the XSS hardening journey, and the production incidents that shaped the design.
The Two Isolated Planes
Shopify Functions (WASM modules that run at checkout) read their configuration from $app: namespace metafields — a private namespace scoped to the app itself, invisible to any other app or to the storefront. Theme Extensions (Liquid templates that render on storefronts) can only read from the app namespace, which is the public metafield namespace exposed to the storefront context — specifically app.metafields.app_config.*.
This creates a fundamental problem:
-
Plane 1 (Checkout): The Discount Function reads rules from
$app:markdo.discount_rules— a JSON metafield containing all rule configurations. This drives what discount actually applies at checkout. -
Plane 2 (Storefront): The 18 Liquid widgets read from
app.metafields.app_config.widget_bindings— a completely separate JSON metafield that contains widget configuration and display data.
If a merchant creates a "20% off orders over $100" rule and enables the Smart Discount Banner widget, the Function needs the rule in Plane 1 to apply the discount, and the widget needs the rule data in Plane 2 to display "20% OFF orders $100+". If Plane 2 is stale or missing, the widget shows nothing — or worse, shows wrong information.
The solution is a sync system that runs on every rule change: the backend reads the canonical rules from Plane 1, builds display-optimized snapshots, and writes them into Plane 2.
The RuleSnapshot Bridge
Every time a merchant creates, edits, or toggles a rule, the sync service runs syncRuleSnapshotsToWidgets(). This function reads all widget bindings from the app_config metafield, finds the matching rule for each binding, and generates a RuleDisplaySnapshot:
function buildRuleSnapshot(rule: ShopifyRuleConfig): RuleDisplaySnapshot {
const snapshot: RuleDisplaySnapshot = {
discountType: rule.discountType,
discountValue: rule.discountValue,
ruleName: rule.name,
ruleActive: rule.isActive,
ruleVersion: rule.revision,
};
switch (rule.discountType) {
case "volume":
if (rule.discountTiers?.length) {
snapshot.tiers = rule.discountTiers.map(t => ({
min: t.minQty, value: t.discount
}));
snapshot.minQuantity = rule.discountTiers[0].minQty;
}
break;
case "spend_save":
if (rule.spendTiers?.length) {
snapshot.tiers = rule.spendTiers.map(t => ({
min: t.minAmount, value: t.discount
}));
snapshot.minPurchase = rule.spendTiers[0].minAmount;
}
break;
case "mix_match":
snapshot.mixMatchConfig = {
minItems: rule.mixMatchConfig.minItems,
discountType: rule.mixMatchConfig.discountType,
discountValue: rule.mixMatchConfig.discountValue,
};
break;
case "bogo":
case "advanced_bogo":
snapshot.bogoConfig = {
buyQty: rule.bogoBuyQty,
getQty: rule.bogoGetQty,
getDiscount: rule.bogoGetDiscount,
};
break;
}
return snapshot;
}
The snapshot is a flattened, display-optimized projection of the rule. It doesn't contain the full rule configuration — just what the storefront needs to render. This matters because metafield values have size limits — most types are capped at 64KB, and JSON metafields at 128KB — and with 18 widgets each potentially bound to a rule, the data needs to be compact.
The Liquid template reads the snapshot through the binding:
{% assign bindings = app.metafields.app_config.widget_bindings.value %}
{% assign binding = bindings['smart-discount-banner'] %}
{% if binding and binding.enabled == true %}
{% assign snap = binding.ruleSnapshot %}
{% assign discount = snap.discountValue | default: 0 %}
{% assign min_purchase = snap.minPurchase | default: 0 %}
<div class="mk-banner" data-discount="{{ discount | escape }}" data-min="{{ min_purchase | escape }}">
{{ 'smart_discount_banner.qualify_pct' | t: value: discount | escape }}
</div>
{% endif %}
Three Metafield Channels
The widget system actually uses three separate metafield channels, each serving a different purpose:
widget_bindings— The primary channel. A JSON object keyed by widget handle (e.g.,"spin-to-win","tier-table"). Each entry containsenabled,settings(visual config like colors),ruleId,ruleActive,ruleVersion, andruleSnapshot. Every widget reads from this.widget_tier_rules— A JSON array of full tier rule data for product-page widgets (tier-table, quantity-discount-badge). These widgets need the complete tier structure with product conditions, not just a snapshot.widget_mix_match— A JSON array of mix-and-match rule configs with eligible collection and product type data. The mix-match-picker widget needs this to filter which products qualify.
Why three channels instead of one? Size limits and access patterns. The bindings metafield is read on every page load by every widget — it needs to be small. The tier rules and mix-match configs are only read by specific widgets on specific pages, so they can afford to be larger.
One prerequisite worth mentioning: for Theme Extensions to read these metafields, the app must define them with storefront access set to PUBLIC_READ in the metafield definition (via the Admin API's MetafieldDefinition input). Without this, the app.metafields.app_config.* namespace is invisible to the storefront context even if the metafield exists.
The i18n Bridge: Liquid to JavaScript
All 18 widgets support 6 languages (English, German, French, Spanish, Japanese, Chinese). Shopify's | t filter handles Liquid-side translation, but JavaScript running in the browser can't call | t. The solution is an "i18n bridge" pattern that serializes all translated strings into a JSON data attribute at render time:
{%- assign t_spin_wheel = 'spin_to_win.spin_wheel' | t -%}
{%- assign t_spin_now = 'spin_to_win.spin_now' | t -%}
{%- assign t_spinning = 'spin_to_win.spinning' | t -%}
{%- assign t_try_again = 'spin_to_win.try_again' | t -%}
{%- assign t_you_won = 'spin_to_win.you_won' | t -%}
{%- assign j_spin_wheel = t_spin_wheel | json -%}
{%- assign j_spin_now = t_spin_now | json -%}
{%- assign j_spinning = t_spinning | json -%}
{%- assign j_try_again = t_try_again | json -%}
{%- assign j_you_won = t_you_won | json -%}
{%- assign i18n_json = '{"spinWheel":' | append: j_spin_wheel | append: ',"spinNow":' | append: j_spin_now | append: ',"spinning":' | append: j_spinning | append: ',"tryAgain":' | append: j_try_again | append: ',"youWon":' | append: j_you_won | append: '}' -%}
<div class="mk-spin-fab"
data-i18n="{{ i18n_json | escape }}">
The key detail is the two-phase approach. Phase 1 applies | json to each translation string individually via intermediate {% assign %} variables. This is critical because in Liquid's pipe chain, | json operates on the accumulated result — writing append: t_spin_wheel | json would JSON-encode the entire concatenated string {"spinWheel":<translation>, not just the translation. Phase 2 assembles the pre-encoded fragments into a valid JSON string, then applies | escape once to HTML-entity-encode the whole thing for safe embedding in an HTML attribute.
Then in JavaScript:
var fab = document.querySelector('.mk-spin-fab');
var t = JSON.parse(fab.dataset.i18n || '{}');
// Later:
button.textContent = t.spinning;
resultEl.textContent = t.youWon;
The | json filter is critical here — it properly escapes strings for JSON embedding, handling quotes, unicode characters, and newlines. But it must be applied to each value before assembly, not during the append chain. The final | escape filter HTML-entity-encodes the complete JSON string, so a French translation containing an apostrophe ("l'offre") or double quotes won't break the attribute boundary. The browser decodes HTML entities before JavaScript reads dataset.i18n, so JSON.parse() receives clean JSON.
This pattern means zero server round-trips for translations. The Liquid template renders once with the customer's locale, bakes all translated strings into the DOM, and JavaScript reads them instantly.
The Currency Symbol Problem
Every widget that displays monetary amounts needs the currency symbol. Shopify does provide {{ localization.currency.symbol }} — the localization object is populated in multi-currency stores based on the buyer's selected currency, but in single-currency stores, or in some App Block rendering contexts, it may be unavailable or return an empty currency. You also can't use it in external JavaScript files where Liquid isn't evaluated.
We maintain a shared snippet (_currency-symbol.liquid) that maps 20 supported currencies as a reliable fallback. It reads localization.currency.iso_code when available, falling back to shop.currency (the store's default):
{% assign active_currency = localization.currency.iso_code | default: shop.currency %}
{% case active_currency %}
{% when 'USD' %}{% assign currency_symbol = '$' %}
{% when 'EUR' %}{% assign currency_symbol = '€' %}
{% when 'GBP' %}{% assign currency_symbol = '£' %}
{% when 'JPY' %}{% assign currency_symbol = '¥' %}
{% comment %} ... 16 more currencies ... {% endcomment %}
{% endcase %}
When localization is unavailable (single-currency store, or certain App Block rendering contexts), the | default: shop.currency fallback ensures we always resolve to something. If localization itself doesn't exist in the rendering context, Liquid treats localization.currency.iso_code as nil and the default filter kicks in — no error, just a silent fallback to shop.currency. The map then gives us a consistent symbol regardless of context.
Widgets originally attempted {% render '_currency-symbol' %} — but {% render %} runs in an isolated scope, so {% assign currency_symbol %} inside the snippet never reaches the parent template. The parent sees currency_symbol as nil. To make a shared snippet work, you'd need a return-style pattern where the snippet outputs the value and the parent captures it:
{% comment %} _currency-symbol.liquid — return-style {% endcomment %}
{% assign active_currency = localization.currency.iso_code | default: shop.currency %}
{% case active_currency %}
{% when 'USD' %}{% assign currency_symbol = '$' %}
{% when 'EUR' %}{% assign currency_symbol = '€' %}
{% when 'GBP' %}{% assign currency_symbol = '£' %}
{% when 'JPY' %}{% assign currency_symbol = '¥' %}
{% comment %} ... 16 more currencies ... {% endcomment %}
{% endcase %}
{{- currency_symbol -}}
The parent template wraps the render in {% capture %} and strips whitespace:
{% capture currency_symbol %}{% render '_currency-symbol' %}{% endcapture %}
{% assign currency_symbol = currency_symbol | strip %}
This works, but adds verbosity and the {% render %} output still introduces newlines that can break inline <script> blocks. Most of our widgets that need the symbol in Liquid templates use the {% capture %} pattern. The exit-intent.liquid widget ended up inlining the entire currency map to sidestep both problems:
{% comment %} Inline currency symbol — avoids whitespace from {% render %} breaking inline JS {% endcomment %}
{% assign active_currency = localization.currency.iso_code | default: shop.currency %}
{% case active_currency %}{% when 'EUR' %}{% assign curr_sym = '€' %}...{% endcase %}
For widgets that use currency in external JavaScript files, the Liquid-side map can't help — the JS needs its own copy. We pass the active currency code via a data-currency attribute (rendered inside the {% if widget_visible %} block, so it only exists when the widget actually renders) so the JS knows which entry to look up:
<div class="mk-widget" data-currency="{{ active_currency | escape }}">
var _cm = {USD:'$',EUR:'€',GBP:'£',JPY:'¥',CNY:'¥',KRW:'₩',INR:'₹',BRL:'R$',
CAD:'C$',AUD:'A$',CHF:'Fr.',SEK:'kr',NOK:'kr',DKK:'kr',PLN:'zł',
MXN:'$',NZD:'NZ$',SGD:'S$',HKD:'HK$',TWD:'NT$'};
var el = document.querySelector('.mk-widget');
var sym = _cm[el.dataset.currency] || '$';
Three parallel implementations of the same currency map — return-style Liquid snippet (via {% capture %}), inline Liquid, and JavaScript — all manually kept in sync. It's not elegant, but it works across all three rendering contexts (Liquid template, inline script, external JS file).
XSS Hardening: 18 Templates, 6 Vulnerabilities
Our security audit found XSS vulnerabilities in 6 of the 18 Liquid templates. The pattern was consistent: user-controlled content (merchant-entered text in block settings) was being injected into the DOM via innerHTML without escaping.
The FOMO notification widget was the most severe case. It displayed customizable notification messages like "John from London just purchased!" — the Liquid template used {{ block.settings.message_text }} for the initial render, and JavaScript for dynamic updates:
<!-- Liquid output without explicit | escape filter -->
<div class="fomo-text">{{ block.settings.message_text }}</div>
// BEFORE (vulnerable)
txEl.innerHTML = d.text; // d.text from merchant config
Shopify's fork of Liquid includes built-in HTML auto-escaping at the template compilation level, so {{ block.settings.message_text }} does get escaped by default. But relying on implicit auto-escaping is fragile — content can be marked safe through various code paths, and the protection isn't always obvious to audit. The safer practice is to always apply an explicit | escape filter as defense-in-depth. The more immediate vulnerability in our case, though, was the JavaScript side: innerHTML interprets any HTML in the string, so a merchant config value containing <img src=x onerror=...> would execute regardless of what Liquid does server-side.
The fix was switching to textContent for all dynamic text updates:
// AFTER (safe)
txEl.textContent = d.text;
The exit-intent widget had a similar issue — it built HTML strings from merchant-configured discount codes and inserted them via innerHTML. The fix was to use DOM APIs (createElement, textContent) instead of string concatenation.
The tier-table widget builds HTML table rows from tier rule data. Since the data comes from parsed JSON (numeric minQty and discount values), the risk is lower — but we still switched to textContent for cell content and only use innerHTML for the static table structure.
The audit also led to extracting all inline JavaScript from Liquid templates into external .js asset files. This serves two purposes: it satisfies Content Security Policy (CSP) requirements that prohibit inline scripts, and it separates concerns — Liquid handles data binding, JavaScript handles interactivity.
Three widgets got their JS extracted: mix-match-picker.liquid, spin-to-win.liquid, and exit-intent.liquid. The pattern is consistent: Liquid renders the DOM structure with data-* attributes containing all necessary data, and the external JS file reads those attributes to initialize behavior.
The Binding Bootstrap: A Three-State Gate
Every one of the 18 widgets starts with an identical "binding bootstrap" — a Liquid preamble that reads the widget's configuration from the shared metafield:
{% assign bindings = app.metafields.app_config.widget_bindings.value %}
{% assign widget_enabled = false %}
{% assign widget_visible = false %}
{% if bindings %}
{% assign binding = bindings['flash-sale'] %}
{% if binding %}
{% if binding.enabled == true %}
{% assign widget_enabled = true %}
{% assign widget_visible = true %}
{% if binding.planLocked == true %}
{% assign widget_visible = false %}
{% endif %}
{% if widget_visible %}
{% assign bs = binding.settings %}
{% if bs.primary_color != blank %}{% assign primary_color = bs.primary_color %}{% endif %}
{% if bs.bg_color != blank %}{% assign bg_color = bs.bg_color %}{% endif %}
{% endif %}
{% endif %}
{% endif %}
{% endif %}
This implements a state gate with four outcomes:
- No binding exists → widget doesn't render at all (merchant hasn't configured it)
-
Binding exists,
enabled: false→ widget is configured but disabled (merchant turned it off) -
Binding exists,
enabled: true,planLocked: true→ widget is enabled but locked behind a paywall (merchant is on a free plan, paid widget unavailable).widget_enabledis true butwidget_visibleis false — the binding data stays intact, so upgrading just flipsplanLockedto false. -
Binding exists,
enabled: true,planLocked: false→ widget renders with the merchant's visual settings
The widget_enabled vs widget_visible distinction matters for the "paid widget on free plan" scenario. When a merchant downgrades, we set planLocked: true on non-free widgets but keep enabled: true and the binding data intact. The widget doesn't render, but if the merchant upgrades again, we flip planLocked to false without needing to rebuild the configuration. In practice, widget_enabled tracks whether the merchant explicitly turned the widget on (useful for logging and debugging), while rendering decisions depend solely on widget_visible.
Production Incidents
Incident 1: Sync Order Catastrophe
A refactor accidentally swapped the order of metafield writes in the sync pipeline. The metaobject write (backup storage) was made to execute first, and when it failed (due to a stale ID), the error short-circuited the entire sync — meaning the Function metafield (Plane 1) was never updated. Automatic discounts silently stopped working for affected merchants.
The fix: restore "metafield first, metaobject as non-blocking backup" ordering, and add self-healing for stale metaobject IDs. The metafield write to Plane 1 happens first; the code below shows the backup metaobject step that follows:
try {
await updateMetaobject(shop, rules);
} catch (e) {
if (isRecordNotFound(e)) {
await clearStaleMetaobjectId(shop);
await createMetaobject(shop, rules); // retry with create path
}
// metaobject is a backup — if it fails, Plane 1 (metafield) already has the data.
// We await it (not fire-and-forget) so the caller can log duration, but the catch
// ensures the failure doesn't propagate to the sync pipeline.
}
Incident 2: Paid Widgets After Downgrade
When merchants downgraded from Pro to Free, the UI correctly showed paid widgets as "locked." But the metafield had no planLocked field set for those widgets (defaulting to falsy), so they continued rendering on storefronts. Merchants were getting paid-tier features for free.
The fix was a lockPaidWidgets() function that runs on plan downgrade, setting planLocked: true on all non-free widgets while keeping enabled and settings intact — so upgrading instantly restores them:
async function lockPaidWidgets(shop: string) {
const bindings = await readBindings(shop);
for (const def of WIDGET_DEFINITIONS) {
if (!def.free && bindings[def.handle]) {
bindings[def.handle].planLocked = true;
}
}
await writeBindings(shop, bindings);
}
Incident 3: Stale Widget Data After Rule Edit
A merchant edited a rule (changed discount from 10% to 20%), but the storefront widget continued showing "10% OFF." The widget bindings stored the ruleId but not the rule version, so there was no way to detect that the rule had changed since the last sync.
The fix was adding ruleVersion (mapped to Rule.revision) to every binding write. The sync writes the new version to both binding.ruleVersion and binding.ruleSnapshot.ruleVersion atomically. The storefront then checks for internal inconsistency:
{% if binding.ruleVersion != binding.ruleSnapshot.ruleVersion %}
{% comment %} Partial sync detected — snapshot wasn't fully updated {% endcomment %}
{% assign widget_visible = false %}
{% endif %}
This catches partial sync failures — where the outer version was bumped but the snapshot wasn't regenerated. It does not catch a complete sync failure where neither field is updated (both remain at the old version, equal to each other). That scenario is handled by the retry and alerting layer described in the Sync Resilience section below: exponential backoff for transient failures, alerting on failure rate spikes, and a manual re-sync button as the escape hatch. The version check is one layer of defense, not the only one.
Incident 4: Priority=0 Treated as Falsy
The tier-rules shared JavaScript used || 999 as a fallback for missing priority values. But priority: 0 is a legal value meaning "highest priority" — and 0 || 999 evaluates to 999, pushing the highest-priority rule to the bottom.
// BEFORE (buggy)
rules.sort((a, b) => (a.priority || 999) - (b.priority || 999));
// AFTER (correct)
rules.sort((a, b) =>
(a.priority != null ? a.priority : 999) - (b.priority != null ? b.priority : 999)
);
Incident 5: Concurrent Metafield Writes
Two simultaneous rule edits from the merchant's admin could trigger concurrent read-modify-write cycles on the same metafield. Thread A reads bindings, Thread B reads bindings, Thread A writes (adding widget X), Thread B writes (adding widget Y) — Thread A's write is lost because Thread B's write was based on the stale pre-X state.
The fix was a per-shop async mutex:
const shopLocks = new Map<string, Promise<void>>();
async function withShopLock<T>(shop: string, fn: () => Promise<T>): Promise<T> {
const prev = shopLocks.get(shop) ?? Promise.resolve();
const next = prev
.catch((e) => {
// Previous operation failed — log but don't cascade the failure.
// Each sync is independent; one merchant's stale data shouldn't block the next write.
console.warn(`[withShopLock] previous op failed for ${shop}:`, e);
})
.then(fn);
const cleanup = next.finally(() => {
if (shopLocks.get(shop) === next) {
shopLocks.delete(shop);
}
});
cleanup.catch(() => {}); // Prevent unhandled rejection from the cleanup chain
shopLocks.set(shop, next);
return next;
}
Each operation chains onto the previous one for the same shop, serializing writes without blocking other shops. The .catch() before .then(fn) is deliberate: if the previous operation failed, we log it and still proceed with the current operation. Each sync is an independent write — one merchant's stale data shouldn't block the next write. The cleanup logic carefully only deletes the Map entry when no successor has chained on, preventing unbounded growth.
Sync Resilience: Retries, Alerts, and Rollbacks
The five incidents above all share a common theme: silent failures. A sync fails, nobody notices, and the storefront shows stale or missing data until a merchant complains. After incident #1 (the sync order catastrophe), we added a resilience layer on top of the sync pipeline:
Structured failure logging. Every sync attempt logs a structured event with shop, ruleId, widgetHandle, attempt, and duration. Failures include the full error object. This lets us query "show me all sync failures for shop X in the last hour" from our log aggregator.
Exponential backoff retry. Transient failures (Shopify API rate limits, network timeouts) retry up to 3 times with exponential backoff (1s, 2s, 4s). Permanent failures (validation errors, missing resources) fail immediately — retrying won't help.
Alerting on failure rate. If a shop's sync failure rate exceeds 10% over a 15-minute window, an alert fires. This catches the "silent degradation" pattern where individual failures are rare but accumulate over time.
Manual re-sync endpoint. The admin panel has a "Re-sync widgets" button per rule that forces a full syncRuleSnapshotsToWidgets() run. This is the escape hatch for when automated sync fails for non-transient reasons (e.g., a bug in the snapshot builder that we've since fixed).
Metafield rollback on corruption. If a sync write produces a metafield that fails JSON parsing on the next read, we log an alert and restore from the metaobject backup (the non-blocking write from incident #1). This is a last resort — the metaobject backup exists precisely for this recovery scenario.
The key insight: sync failures are not "if" but "when." Shopify's API has rate limits, network partitions happen, and metafield size limits surprise you. The system needs to detect failures, retry what's retryable, alert on what's not, and provide manual override for edge cases.
The Shared Module Pattern
Some widgets share logic. The tier-table and quantity-discount-badge widgets both need to resolve tier rules from a shared pool, match them to the current product, and sort by priority. Instead of duplicating this logic, we extracted it into tier-rules-shared.js:
window.MkTierRules = {
resolveTiers: function(raw) { /* parse + normalize tier rules */ },
matchesProduct: function(rule, productId, productType, tags, collections) {
/* check product conditions */
}
};
Both widgets load the shared script via <script src="{{ 'tier-rules-shared.js' | asset_url }}" defer>, then use a retry-based initialization to handle the race condition where the deferred script hasn't loaded yet:
var retries = 0;
function init() {
if (!window.MkTierRules) {
if (retries++ < 20) { setTimeout(init, 50); }
return;
}
var tiers = window.MkTierRules.resolveTiers(rawRules);
// ... proceed with rendering
}
init();
20 retries at 50ms intervals gives a 1-second window for the shared script to load — generous enough for slow connections, short enough to not noticeably delay rendering. The rawRules variable is read from a data-tier-rules JSON attribute on the widget root element, which the Liquid template renders from the widget_tier_rules metafield channel.
What I Learned
After shipping 18 widgets across 6 languages with a dual-plane sync architecture, here's what I'd tell anyone building a similar system:
Design the sync boundary first. The split between Plane 1 (checkout Function data) and Plane 2 (storefront widget data) is the hardest part to get right. Get the snapshot schema wrong, and you'll either blow the metafield size limit or leave widgets without the data they need.
Version everything — but don't rely on versions alone. Rule revisions, binding versions, snapshot versions — version numbers let you detect partial sync failures where one layer was updated but another wasn't. But they can't detect a complete sync failure where nothing was updated. For that, you need retry logic, failure alerting, and a manual re-sync escape hatch.
The i18n bridge pattern scales well. Serializing translations into
data-i18nJSON attributes is a bit verbose, but it's mechanical, auditable, and has zero runtime cost. We added Japanese and Chinese late in the project with zero JavaScript changes.Inline JavaScript is technical debt. CSP compliance, maintainability, and browser caching all push toward external JS files. Extract early — retrofitting is painful when you have 18 templates with inline scripts.
Metafield concurrency is real. Any app that does read-modify-write on metafields needs serialization. The per-shop mutex pattern is simple and effective, but you need to think about cleanup to prevent memory leaks.
Markdo is a Shopify discount app with 18 storefront widgets and a Rust-based checkout Function. The widget system described here is part of the Theme Extension that renders discount information on merchant storefronts.
Top comments (0)