By Zhonglian Network — custom software development, Nanning, China
Most architecture writing is about greenfield systems. That's the easy part. The hard question is: what makes a system still running, unchanged in its core, after a decade?
We have three systems that have crossed or are crossing that line:
- A school cafeteria procurement platform serving 140+ schools — in production for over 10 years, never rewritten.
- A POS payment system built on the UnionPay 8583 protocol — still the core payment architecture for multiple commercial deployments.
- An education platform serving 100+ institutions and 1M+ end users, holding roughly 15 TB of data.
None of these are clever. That's the point. Here's what actually mattered.
1. The database schema is the real API
Application code gets rewritten every few years. The schema doesn't. Every long-lived system we have owes its survival to schema decisions made in week one.
What worked:
-
Explicit columns over EAV. Early on, someone always proposes an entity-attribute-value table "so we can add fields without migrations." We've seen this fail every time. A procurement system needs
purchase_order.amountto be aDECIMAL(18,2)with a constraint — not a row in an attributes table. Query performance and data integrity both collapse under EAV. - Never delete. Add a status. Financial records get soft-deleted with an audit trail. The school procurement system can reconstruct every state a purchase order passed through, ten years back. That's not a feature we built later — it was designed in.
- Surrogate keys everywhere. Natural keys (school code, student number) change. We learned this the expensive way: a primary key that changes is a primary key you'll regret.
- Money as integer minor units or DECIMAL, never FLOAT. Obvious, and still violated constantly.
What we'd do differently: we normalized aggressively in 2012. Some of those joins are now the slowest part of the system. A few denormalized read columns for reporting would have paid for themselves many times over.
2. Financial systems need idempotency at the protocol layer
The UnionPay 8583 POS system is the clearest example of a decision that has aged well.
8583 is a byte-oriented financial messaging protocol — fixed-format bitmaps and fields, sent over TCP. The critical property is that the network is not reliable, but the transaction must be exactly-once.
Our design, simplified:
POS terminal Our gateway UnionPay
| | |
|-- 0200 (sale req) ------>| |
| |-- 0200 ---------------->|
| |<-- 0210 (response) -----|
|<-- 0210 (response) ------| |
| | |
[timeout / no response] | |
|-- 0200 (retry) --------->| |
| [dedup by STAN + RRN] |
Three rules that made this work:
- Every request carries a unique trace number (STAN) plus a retrieval reference number (RRN). The gateway keeps a dedup table keyed on those. A retry with the same key returns the original result — it does not execute a second charge.
- The gateway, not the terminal, is the source of truth. Terminals cache but never decide. If a terminal is destroyed, no money is lost.
- Reconciliation runs daily and is non-optional. A batch job compares our ledger against UnionPay's settlement file. Any mismatch is an incident, not a warning. In ten years, this has caught issues that would otherwise have become customer-visible.
The generalizable lesson: for anything involving money, design the retry path before the happy path. Ask "what happens if this message arrives twice?" for every write endpoint. If the answer requires human intervention, the design is wrong.
3. Multi-tenancy: separate the data, share the schema
The 140-school procurement platform is multi-tenant. We chose a shared schema with a tenant_id column — not separate databases per school.
The trade-off:
| Approach | Pros | Cons |
|---|---|---|
| Shared schema + tenant_id | One migration path, cheap ops, easy cross-tenant reporting | Risk of tenant data leakage if a query forgets the filter |
| DB per tenant | Strong isolation, per-tenant backup/restore | 140 migrations per release; ops cost scales linearly |
We chose shared schema. To make it safe, we enforced tenant filtering at the data-access layer, not in hand-written SQL — every query goes through a repository that injects the tenant predicate. Hand-written SQL is where isolation bugs live.
If we were doing it today with stricter compliance requirements, we'd likely use row-level security in the database itself rather than trusting the application layer.
4. Scaling to 1M users and 15 TB: boring beats clever
The education platform taught us that scale problems are usually access-pattern problems, not capacity problems.
What actually mattered:
-
Pagination discipline. Never
SELECT *without a bounded range. The first version of our admin console loaded a full institution list; that page became unusable around 200k rows. Keyset pagination fixed it permanently. - Move files out of the database early. Binary blobs in the DB made backups grow linearly with usage. Object storage plus a reference table changed our backup window from hours to minutes.
- Read replicas for reporting. Reports were competing with transactional traffic for the same locks. Isolating them was a one-week change that removed an entire class of incidents.
- Partition the biggest tables by time. The 15 TB is dominated by a handful of append-only tables. Time-based partitioning made archival and retention possible.
None of this is novel. It's just that teams skip it until it hurts.
5. What we'd tell a team starting today
Five rules, ordered by how much pain they prevent:
- Write the schema as if you'll still be querying it in 2036. Because you will.
- Design the idempotency key before the endpoint. Especially for payments, orders, and anything that sends a message.
- Enforce invariants in the database, not only in the service layer. Application code has bugs; constraints don't.
- Keep an audit trail from day one. Retrofitting one is a project; designing one is a column.
- Prefer the boring technology. The systems that survive are the ones whose dependencies still exist. We still run .NET and Oracle/MySQL systems from 2012 without drama.
A note on what this means for offshoring
If you're evaluating an offshore partner, the question that separates vendors isn't "what's your stack?" — it's "show me something you built that's still running."
Anyone can ship an MVP. Ask for the system that's been in production for eight years, and ask what they'd change. A team that has maintained long-lived systems will answer that question with specifics, not adjectives.
About Zhonglian Network
Software development company based in Nanning, China, operating since 2012.
- Custom software development — web, mobile, desktop
- Smart ticketing and access-control systems
- POS / membership / payment systems (UnionPay 8583 protocol)
- AI application development and enterprise system integration
Stack: .NET Core / C#, Python, Go, Java, Oracle, MySQL, Redis, Vue, React, Flutter, Taro
Track record: 100+ institutions and 1M+ end users on one platform; a procurement system in production 10+ years; ticketing systems deployed commercially across multiple provinces. 5 national invention patents.
Contact: xiaoru@zl771.cn
Tags: software architecture, legacy systems, payment systems, UnionPay 8583, offshore development
Top comments (0)