Many multi-brand QSR and retail estates span clouds in a recognisable way. Staff sign in with Microsoft Entra ID, some back-office systems run on SQL Server in Azure, the data platform is on Google Cloud, and a loyalty vendor or an acquired brand brings AWS. If your estate looks like that, it probably isn't anyone's plan. Each piece was a sensible decision when it was made. The cost shows up later, in the connections between them: service account keys copied into another cloud's secret store, the same customer held under three different IDs, and nightly copies of the same tables in two places.
Our earlier post Multi-Cloud Versus Consolidation made the cross-industry case that most organisations are multi-cloud whether or not they chose to be, and that the work is to make the estate coherent. This post applies that to retail and QSR. It looks at how these estates come to span three clouds, where the extra engineering falls, and the identity and data layer that holds them together, with a diagram of the layer, the token exchange behind it, and code for the customer key at its centre.
Joining one customer across three clouds
Take an estate where the catering and gift card systems run on SQL Server in Azure, and a back-office sync job moves their changes into BigQuery on Google Cloud, which is the data platform for sales, operations, and marketing. The loyalty programme runs on a SaaS platform, with an integration in the retailer's AWS account that syncs member records, and a history export that the vendor writes to an S3 bucket in the same account. Delivery orders arrive as webhooks from each aggregator at an integration service on Google Cloud, which injects them into the POS and forwards them to the data platform. In our experience, what aggregators pass on about the customer varies by aggregator and by market, and some orders carry no usable email at all, while others carry a masked relay address that won't match anything else. Marketing wants to send a customer an offer, so the data platform needs to know that a catering order, a loyalty member, and a delivery order with the same email address belong to the same person, and to treat delivery orders without one as unmatched, flagging relay-address keys so they don't inflate customer counts.
Three teams own the three parts of that question, and each has solved its own part well. The back office stores the email typed in at checkout. The loyalty platform has a member ID. The delivery integration stores the aggregator's customer email, when the order carries one. Joining them means moving email addresses, or full customer tables, between clouds, and each team has a reasonable fix that makes the overall picture a little harder to manage.
How estates end up on three clouds
Most of the clouds in an estate like this arrive with a product or a business decision. Microsoft 365 often brings Entra ID with it, and Entra ID becomes the place staff identities live. Back-office applications built on SQL Server often move to Azure, a short step from the data centre they ran in. A data team picks Google Cloud for the data platform on technical grounds. A SaaS loyalty platform runs on whichever cloud its vendor chose, and its exports tend to land nearby. An acquired brand brings its own cloud accounts, its own POS integration, and a team that knows them.
Franchise networks add another layer, because franchisees may run systems the franchisor doesn't control and send data in whatever form the franchise agreement provides for.
Most of the cost comes from the default way of connecting the pieces. A workload on one cloud that needs something on another usually gets a long-lived credential, such as a service account key or an access key, stored in its own cloud's secret manager. Each team identifies customers its own way, so matching them means exporting email addresses or customer tables to wherever the match runs. And data needed in two clouds tends to get copied on a schedule, so there are two copies to secure, two to keep current, and a transfer bill each time it moves.
Sometimes the right answer is consolidation. Sakura Sky moved Craveable Brands' data platform to Google Cloud, bringing data from across its brands and outlets into one BigQuery warehouse. Where consolidation doesn't pay, because the loyalty vendor's cloud isn't yours to choose or the SQL Server estate works well where it is, the job is to make the connections between clouds deliberate.
Where the cross-cloud work falls
For a retail estate, most of that work falls in three places.
The first is machine identity. Services in one cloud need to call services and read data in another without anyone storing a long-lived key for the other cloud. AWS, Google Cloud, and Microsoft Entra ID all document ways to exchange one platform's credentials for another's short-lived tokens (AWS, n.d.a; Google Cloud, n.d.a; Microsoft, 2025), so the stored key can usually be retired.
The second is customer identity. Every system that holds customer records needs to agree on who a customer is, and it helps if they can do that without each one holding the others' raw contact details. That usually means one pseudonymous customer key, computed by one service.
The third is data placement. Each dataset of record lives in one place, with one owner, and other clouds read it where it is. Copies are made for a reason, such as latency, data residency, or resilience, and each copy is recorded with its source and refresh schedule.
Two smaller disciplines sit alongside them. Events crossing clouds, such as aggregator orders, use one envelope format, so each consumer doesn't write its own parser for each source. And consent travels with the customer key: an opt-out recorded in the loyalty platform has to reach the data platform before the next campaign runs, which Part 1 covered from the loyalty side.
The identity and data layer
Figure 1 shows one way the pieces fit together for the estate above, with Google Cloud as the data platform and the home of the customer key service. In your estate the clouds may play different roles, with AWS as the data platform, for example, or the loyalty platform on Azure, and the same pattern should carry over, though the services and their limits differ.
Machine identity with fewer stored keys
Google Cloud's Workload Identity Federation lets workloads on Azure and AWS exchange their own environment's credentials for short-lived Google Cloud credentials, so they can call Google Cloud APIs without a service account key (Google Cloud, n.d.a). Figure 2 shows the Azure case. The back-office sync job runs on an Azure VM with a managed identity. It asks the Azure Instance Metadata Service for an access token for a Microsoft Entra ID application set up for federation, identified by its Application ID URI. In Google Cloud, a provider in a workload identity pool takes the tenant's issuer as its issuer URL and that Application ID URI as its allowed audience. Google's example issuer is https://sts.windows.net/<tenant-id>, but the value can vary, so it's copied from the iss claim of a real token. The token's subject is the managed identity's object ID, which the provider maps to the federated identity's subject, so IAM roles can be granted to that one identity, and an attribute condition on the provider can reject tokens from any other identity in the tenant (Google Cloud, n.d.a). Google Cloud's Security Token Service then exchanges the Entra token for a short-lived Google Cloud access token, referred to here as a federated token. BigQuery's API has no known limitations with federated identities (Google Cloud, n.d.g), so the sync job can be granted access to its datasets directly.
The AWS side of Figure 1 works the same way with different credentials. One of the options Google documents for AWS uses the workload's ordinary temporary AWS credentials, which Google Cloud verifies through AWS's GetCallerIdentity API with no configuration changes in the AWS account. The other uses AWS outbound identity federation, where AWS acts as the identity provider and its workloads request short-lived JSON Web Tokens to present to external services such as Google Cloud (AWS, n.d.b; Google Cloud, n.d.a). Google recommends granting the federated identity access to resources directly, and falling back to service account impersonation only for products that don't support that (Google Cloud, n.d.a).
Calls in the other direction follow the same pattern. Entra ID's workload identity federation lets an app registration or user-assigned managed identity trust tokens from an external provider, with Google Cloud and AWS among the documented scenarios (Microsoft, 2025). AWS's AssumeRoleWithWebIdentity returns temporary AWS credentials, by default for one hour, in exchange for a token from an OpenID Connect provider such as Google, provided the role's trust policy names that provider (AWS, n.d.a). We'd also restrict the subject and the audience in the trust policy, because naming only the provider trusts every identity it issues tokens for. For Google tokens there's a trap: the accounts.google.com:aud condition key takes the token's azp claim when that claim is set, and the token's aud is then checked with accounts.google.com:oaud instead (AWS, n.d.c).
What changes in practice is where trust is configured. Instead of a key in a secret store, each cloud holds a trust relationship naming the other cloud's identity provider, plus conditions on which workloads it accepts. Those relationships need the same care a key did. Revoking access happens on the receiving side, because a token that's already been issued typically stays valid until it expires. In Google Cloud, that means removing the role granted on the resource itself, such as access to the BigQuery dataset or the Cloud Run invoker role held by the impersonated service account, because IAM checks it on each call. Disabling the provider stops new token exchanges, but tokens already issued run until they expire. In AWS, revoking the role's active sessions denies credentials issued before that point (AWS, n.d.d). A provider that accepts any token from an Entra tenant, or any identity in an AWS account, trusts far more workloads than the ones that need access, so its attribute condition names only the specific managed identities or AWS roles that do.
For people, the Entra ID that head-office staff already use can also sign analysts and engineers in to Google Cloud through Workforce Identity Federation, without copying their accounts into Google (Google Cloud, n.d.d), though not every Google Cloud product supports federated identities in the same way (Google Cloud, n.d.g).
One customer key, computed by one service
The customer key in Figure 1 is a pseudonymous join key: a value derived from a customer's normalised email address that every system can store and join on in place of the address itself. Listing 1 computes it with HMAC-SHA256 and a secret key. A plain hash such as SHA-256 of the email isn't enough, because anyone with a list of email addresses can hash them and compare. With HMAC, customer keys can only be produced by whoever holds the secret key.
That's also why the key service runs in one place. If the secret key were copied into every cloud so each could compute customer keys locally, it would become exactly the kind of shared secret federation removes. Instead, the back-office sync job and the loyalty integration authenticate to the key service through federation, the delivery integration calls it with its own Google Cloud identity, and each sends an email address and stores the customer key it returns. If the key service runs on Cloud Run, the federated callers can't be granted access directly, because Cloud Run doesn't support direct resource access for federated identities (Google Cloud, n.d.g). Instead, they use their federated token to impersonate a service account that may invoke the service, and request an ID token as that account, which Cloud Run requires (Google Cloud, n.d.a). Addresses still cross clouds, but only to one authenticated endpoint that doesn't store or log them, in place of whole customer tables copied to wherever the match runs.
The secret itself can live in Cloud KMS, which supports HMAC keys for MAC signing (Google Cloud, n.d.e), so the service calls KMS to compute each HMAC; Listing 1 takes the key as a parameter to keep the example self-contained. Any workload allowed to call the service can still check whether an address matches a stored customer key, so access is limited to named workloads, rate-limited, and logged without the addresses. For the same reason, only the key service's own identity should hold permission to sign and verify with the KMS key, because anyone who can compute or check the HMAC can test addresses against stored keys. Each key also costs a network round trip and a KMS request, so order and enrolment flows don't wait on it: a record can be stored first and keyed asynchronously, and bulk jobs such as backfills and re-keying go through a batch endpoint, where the extra time and cost are easiest to absorb. Campaign audiences leave the data platform as customer keys, and the system that holds the address, such as the loyalty platform or the email service, resolves them for sending.
Python
import hashlib
import hmac
import unicodedata
_EDGE_WHITESPACE = " \t\r\n" # trimmed identically in every implementation
def _ascii_lower(s: str) -> str:
# bytes.lower() changes only A-Z, so non-ASCII characters are left as-is
return s.encode("utf-8").lower().decode("utf-8")
def normalise_email(raw: str) -> str:
s = unicodedata.normalize("NFC", raw.strip(_EDGE_WHITESPACE))
local, at, domain = s.rpartition("@")
if not at or not local or not domain:
raise ValueError("not an email address")
# Policy choice: lowercase the local part too (see RFC 5321, section 2.4)
return f"{_ascii_lower(local)}@{_ascii_lower(domain)}"
def customer_key(raw_email: str, key: bytes, key_version: str) -> str:
"""Pseudonymous join key: the same email always gives the same key."""
if len(key) < 32:
raise ValueError("key must be at least 32 bytes")
message = normalise_email(raw_email).encode("utf-8")
digest = hmac.new(key, message, hashlib.sha256).hexdigest()
return f"{key_version}:{digest}"
Rust
use hmac::{Hmac, Mac};
use sha2::Sha256;
use unicode_normalization::UnicodeNormalization;
type HmacSha256 = Hmac<Sha256>;
// trimmed identically in every implementation
fn is_edge_whitespace(c: char) -> bool {
matches!(c, ' ' | '\t' | '\r' | '\n')
}
pub fn normalise_email(raw: &str) -> Result<String, &'static str> {
let s: String = raw.trim_matches(is_edge_whitespace).nfc().collect();
let (local, domain) = s.rsplit_once('@').ok_or("not an email address")?;
if local.is_empty() || domain.is_empty() {
return Err("not an email address");
}
// Policy choice: lowercase the local part too (see RFC 5321, section 2.4)
Ok(format!("{}@{}", local.to_ascii_lowercase(), domain.to_ascii_lowercase()))
}
/// Pseudonymous join key: the same email always gives the same key.
pub fn customer_key(raw_email: &str, key: &[u8], key_version: &str) -> Result<String, &'static str> {
if key.len() < 32 {
return Err("key must be at least 32 bytes");
}
let message = normalise_email(raw_email)?;
let mut mac = HmacSha256::new_from_slice(key).map_err(|_| "invalid key")?;
mac.update(message.as_bytes());
let digest: String = mac
.finalize()
.into_bytes()
.iter()
.map(|b| format!("{b:02x}"))
.collect();
Ok(format!("{key_version}:{digest}"))
}
Listing 1. A pseudonymous customer key from a normalised email address. Both implementations must return the same bytes for the same input.
Most of the listing is normalisation, and that's where the hard decisions are. The key only works as a join if every implementation produces the same bytes for the same address. That can mean two languages even inside one service, for example when the service is rewritten, or when its batch endpoint runs in a different language from the online one, and the obvious library calls don't always agree. Python's str.strip() with no arguments, for example, removes a few control characters that Rust's trim() keeps. So the listing trims a fixed set of characters, applies Unicode NFC normalisation so that "é" typed as one character or as "e" plus an accent gives the same key, and lowercases only ASCII letters. Both implementations are tested against the same set of input and output vectors, with the library versions they were generated on pinned.
Lowercasing the local part, the bit before the @, is a policy decision. RFC 5321 says the local part must be treated as case-sensitive, while discouraging anyone from relying on that (Klensin, 2008). Lowercasing it accepts a small chance of merging two people's addresses in exchange for fewer split records. Leaving non-ASCII letters alone means "ÉLODIE" and "élodie" produce different keys, which is the cost of keeping the result identical in every language.
The version prefix on each key makes rotation possible. A new secret key gets a new version. Re-keying means recomputing from the email address, so it can only run where addresses are still held, such as the back office or the loyalty platform. Systems that hold only keys, such as the data platform, are re-keyed from an old-to-new key table produced where the addresses are, and that table should be held as tightly as the secret and deleted once re-keying finishes, because anyone holding it and the old secret can link new keys to addresses. Keeping an address-to-key mapping in the key service would make re-keying easier, but it turns the service into a central store of email addresses that needs the same protection as the secret. Because of that cost, some teams rotate only when a key may have been exposed. During a rotation both versions stay live, and the service issues both for new records until every stored record has been re-keyed.
A customer key isn't anonymous. Under the GDPR, pseudonymisation means data can't be attributed to a person without additional information that is kept separately and protected (European Parliament and Council, 2016, Art. 4(5)), and pseudonymised data that could be attributed using that information is still personal data (European Parliament and Council, 2016, Recital 26). California's CCPA defines pseudonymization along similar lines (California Legislature, n.d., §1798.140(aa)). The key cuts the number of systems that need raw addresses to join records, and makes erasure and access requests easier to trace, but the same privacy obligations apply to every system that stores it.
Data read where it lives
In Figure 1, the loyalty vendor's history export lands in an S3 bucket in the retailer's own AWS account, and the loyalty integration writes a member-to-customer-key map beside it. Keeping the map beside the export lets BigQuery Omni, covered in the next paragraph, join loyalty history to customer keys inside AWS, so only keyed, filtered results cross to Google Cloud. Vendor exports are often CSV or Parquet files. If the vendor writes Apache Iceberg tables, or the retailer converts the export, engines such as Spark, Trino, and Flink can work with the same tables at the same time (Apache Iceberg, n.d.).
BigQuery Omni is one way for the data platform on Google Cloud to use that data without copying it first. It runs the BigQuery query engine in AWS and Azure regions over data in S3 or Azure Blob Storage, through BigLake tables, so processing happens where the data sits (Google Cloud, n.d.b). Results can come back to Google Cloud or be written straight to S3. The options differ in practice: cross-cloud joins suit one-off queries over data that isn't large, with a limit of 60 GB per transfer, while materialised view replicas suit repeated queries such as a dashboard. Data moved out of AWS this way still incurs AWS egress charges (Google Cloud, n.d.b), so a filtered read, such as one store-day of loyalty redemptions for one brand, should cost less than moving the whole table. For Iceberg tables in S3 that are managed by a catalog such as AWS Glue, Google now points to cross-cloud data access through its Lakehouse runtime catalog instead (Google Cloud, n.d.c), so the right route depends on how the export is catalogued.
The SQL Server systems are a different case. Their data changes constantly and the data platform needs it alongside sales from every other source, so the sync job reads SQL Server's change data capture and merges changed rows into BigQuery. A managed service can do the same job. Google's Datastream reads change tables or transaction logs from self-managed SQL Server, including SQL Server on a cloud VM, and change tables from Azure SQL Database at service objective S3 (a Standard-tier size, unrelated to Amazon S3) or above (Google Cloud, n.d.f), and can stream them to BigQuery (Google Cloud, n.d.h). For SQL Server it doesn't support Windows Active Directory authentication (Google Cloud, n.d.f), and its setup creates a SQL Server login with a password for the connection (Google Cloud, n.d.h), so there's a stored credential to manage, which is worth weighing against a sync job that uses its managed identity. Replicating the email column as it is in the source means email addresses arrive in BigQuery unkeyed and need keying and removal there, which brings raw addresses back into the data platform. Excluding the email column avoids that, but then those rows have to be keyed at source. Either way it's a deliberate copy, with a known source, owner, and cadence.
A copy is worth making when it has a specific reason. A dashboard that refreshes every few minutes, a dataset that has to sit in a particular country, a replica kept for recovery, or operational data that has to be joined with sales and other sources, like the SQL Server feed, can each justify one. Each of those copies goes in the same record of copies, so that it can be found, refreshed, and eventually removed.
Events need the same treatment at a smaller scale. Aggregator webhooks arrive in each aggregator's own format. CloudEvents, a CNCF graduated specification for describing event data in a common way, gives every order event the same envelope, and its adopters include Amazon EventBridge, Azure Event Grid, and Google Cloud Eventarc (CloudEvents, n.d.). Wrapping each aggregator's order in that envelope at the integration service means downstream consumers in any cloud read one format, and once the integration service has a customer key for an order from the key service, the event can carry the key in place of the email address.
What the estate looks like with this in place
A few things change that teams outside the platform group will notice. Head-office staff use their Entra ID identity in Azure and Google Cloud. Long-lived credentials between clouds are the exception, kept only where a service such as Datastream, or a vendor integration with the S3 bucket that only supports access keys, requires one, and then recorded and rotated; elsewhere each cloud holds trust relationships with conditions narrowed to the workloads that need them. Every system that stores customers stores the same customer key, issued by one service, and an opt-out or an erasure request can be traced across all of them by that key. Each dataset of record has one owner and one location, and every copy has a recorded reason. Aggregator events reach every consumer in one envelope.
The estate also becomes easier to change. Moving the back office off SQL Server, replacing the loyalty platform, or bringing an acquired brand's data in becomes mostly a matter of re-pointing trust relationships and readers, while the customer key and event format can stay the same. Deciding which workloads stay where they are, and which are worth consolidating, is placement work Sakura Sky's Cloud practice does with retail and QSR teams.
Trust conditions and key versions drift as vendors and teams change, and reviewing them is ongoing work Sakura Sky's Managed Services team takes on.
References
Apache Iceberg, n.d. Apache Iceberg™: The open table format for analytic datasets. Apache Software Foundation. Available at: https://iceberg.apache.org/ [Accessed 9 October 2026].
AWS, n.d.a. AssumeRoleWithWebIdentity. AWS Security Token Service API Reference. Available at: https://docs.aws.amazon.com/STS/latest/APIReference/API_AssumeRoleWithWebIdentity.html [Accessed 9 October 2026].
AWS, n.d.b. Getting started with outbound identity federation. AWS Identity and Access Management User Guide. Available at: https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_providers_outbound_getting_started.html [Accessed 9 October 2026].
AWS, n.d.c. IAM and AWS STS condition context keys. AWS Identity and Access Management User Guide. Available at: https://docs.aws.amazon.com/IAM/latest/UserGuide/reference_policies_iam-condition-keys.html [Accessed 9 October 2026].
AWS, n.d.d. Revoke IAM role temporary security credentials. AWS Identity and Access Management User Guide. Available at: https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_use_revoke-sessions.html [Accessed 9 October 2026].
California Legislature, n.d. California Civil Code §1798.140 (California Consumer Privacy Act). California Legislative Information. Available at: https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140 [Accessed 9 October 2026].
CloudEvents, n.d. CloudEvents: A specification for describing event data in a common way. Cloud Native Computing Foundation. Available at: https://cloudevents.io/ [Accessed 9 October 2026].
European Parliament and Council, 2016. Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data (General Data Protection Regulation). Official Journal of the European Union, L 119, 4 May, pp. 1-88. Available at: https://eur-lex.europa.eu/eli/reg/2016/679/oj [Accessed 9 October 2026].
Google Cloud, n.d.a. Configure Workload Identity Federation with AWS or Azure VMs. Identity and Access Management documentation. Available at: https://docs.cloud.google.com/iam/docs/workload-identity-federation-with-other-clouds [Accessed 9 October 2026].
Google Cloud, n.d.b. Introduction to BigQuery Omni. BigQuery documentation. Available at: https://docs.cloud.google.com/bigquery/docs/omni-introduction [Accessed 9 October 2026].
Google Cloud, n.d.c. Create Apache Iceberg external tables. BigQuery documentation. Available at: https://docs.cloud.google.com/bigquery/docs/iceberg-external-tables [Accessed 9 October 2026].
Google Cloud, n.d.d. Workforce Identity Federation. Identity and Access Management documentation. Available at: https://docs.cloud.google.com/iam/docs/workforce-identity-federation [Accessed 9 October 2026].
Google Cloud, n.d.e. MAC signatures. Cloud Key Management Service documentation. Available at: https://docs.cloud.google.com/kms/docs/mac-signatures [Accessed 9 October 2026].
Google Cloud, n.d.f. Stream data from SQL Server databases. Datastream documentation. Available at: https://docs.cloud.google.com/datastream/docs/sources-sqlserver [Accessed 9 October 2026].
Google Cloud, n.d.g. Identity federation: products and limitations. Identity and Access Management documentation. Available at: https://docs.cloud.google.com/iam/docs/federated-identity-supported-services [Accessed 9 October 2026].
Google Cloud, n.d.h. Configure a self-managed SQL Server database for CDC. Datastream documentation. Available at: https://docs.cloud.google.com/datastream/docs/configure-self-managed-sqlserver [Accessed 9 October 2026].
Klensin, J., 2008. RFC 5321: Simple Mail Transfer Protocol. Internet Engineering Task Force. Available at: https://datatracker.ietf.org/doc/html/rfc5321 [Accessed 9 October 2026].
Microsoft, 2025. Workload identity federation concepts. Microsoft Entra Workload ID documentation. Available at: https://learn.microsoft.com/en-us/entra/workload-id/workload-identity-federation [Accessed 9 October 2026].


Top comments (0)