DEV Community

Umar Bilal
Umar Bilal

Posted on Originally published at muhammadumar89.github.io

System design for physical AI: a fully air-gapped stack for an oil and gas operation in Pakistan

This is the engineering summary of a full reference architecture. The paper, its object model as JSON and the model register are open: https://muhammadumar89.github.io/codeninja-research/sovereign-hse-pakistan/.

An oil and gas operator in Pakistan wants one thing from AI in its health, safety and environment department: a warning before the next incident, not a report after it. The data to do that already exists, spread across SAP EHS, SCADA and fire-and-gas historians, camera feeds and scanned investigation files. The constraint is just as clear. None of it may leave the operator's own infrastructure, and no third-party AI API may sit in the serving path.

Here is how that constraint turns into hardware, models and money.

1. Pick models by licence first

On an air-gapped platform you cannot call a hosted model, so every model must be self-hosted, and the licence decides whether the operator owns what it runs. Every pick lets the operator hold, run and fine-tune the weights inside its own boundary:

Role Model Licence
Reasoning, cited answers, agents GLM 5.3 open weights, 753B mixture-of-experts at FP8 bespoke; purely internal use is exempt from its managed-service review
Time-series anomaly and early warning amazon/chronos-2 Apache-2.0
Vision detection baseline Roboflow/rf-detr-large (Nano to Large only) Apache-2.0
Tracking across frames Roboflow trackers Apache-2.0
Multilingual retrieval (English, Urdu, Roman Urdu) BAAI/bge-m3 MIT
OCR of scanned permits and reports PaddlePaddle/PaddleOCR-VL-1.6 Apache-2.0

RF-DETR's larger checkpoints ship under a different platform licence, so the design stops at Large.

2. Size the central tier from the weights, not the brochure

Take the largest filed parameter count, multiply by bytes per parameter at the serving precision, then add a planning factor for the KV cache and activations so long incident histories fit:

weights   = parameters x bytes per parameter
need      = weights x 1.2 planning factor
nodes     = ceil(need / (cards per node x memory per card))
Enter fullscreen mode Exit fullscreen mode

GLM 5.3 is filed at 753B parameters. At FP8 that is 753 GB of weights and 904 GB with headroom, so one node of eight 141 GB cards (1,128 GB) holds it, leaving 375 GB beside the weights for KV cache: long incident histories and concurrent users.

3. Keep fast things at the edge

Detection and forecasting must keep up with cameras and sensors even if the link to the central tier drops. Detection runs on edge nodes inside the plant network, reusing the operator's NPU compute where it exists. Forecasting, OCR and embeddings run on a site inference server beside the historian: Chronos-2, PaddleOCR-VL 1.6 and BGE-M3 together weigh under 4 GB. Edge and site compute are sized by stream and decode load, not by model count.

4. One clock, one backbone, read-only adapters

  • Every source enters through an adapter that only reads, tags provenance, maps to the object model once, is replayable and degrades honestly.
  • Apache Kafka on KRaft orders events per equipment key, so a developing event is read in the order it happened.
  • Chrony with a GNSS grandmaster gives sensors, cameras and servers one clock. Correlating SCADA with camera detections is meaningless if the timestamps drift.

5. What it costs, against the cloud

Three years, public prices, electricity at Pakistan's B3 industrial tariff:

Option Three-year cost (USD)
Own the hardware, with support and power about 670,000
Rent the same GPUs, AWS UAE region, deepest three-year plan (EC2 Instance Savings Plan, paid up front) 1.14 million
Rent the same GPUs, AWS UAE region, on demand 2.82 million
Buy a closed frontier model by the token, 50 users 1.04 to 2.95 million

No hyperscaler runs a region inside Pakistan, so every rented option also moves the data abroad. The full workings and sources are in Appendix A of the paper (version 2, 3 October 2026: the three-year row now uses AWS's deepest plan; version 1 used a 26 percent plan and showed 1.85 million).

6. The one hard dependency

141 GB HBM-class accelerators need a US export licence for Pakistan (Country Group D:4). Approved channels have delivered thousands of GPUs to Pakistani operators, and the design's first phase confirms the installed inventory before anything is bought, with a fallback to a mid-size model on existing hardware.

Reuse it

The object model ships as JSON in the repository in a format meant for import into an ontology platform, with every object's anchor system, properties, status vocabulary and links. Take it, change it, cite it.

Umar Bilal, Cofounder of CodeNinja. CodeNinja is a Middle Eastern American artificial intelligence lab that puts autonomy in physical operations.


Designed on Praxis, CodeNinja's platform for designing physical AI systems. The object model imports into Hyper Ontology, which turns it into a living system. Load it yourself with the open hyper-ontology loader.

Top comments (0)