DEV Community

Umar Bilal
Umar Bilal

Posted on Originally published at codeatoms.ai

System design for physical AI: forecasting construction schedule slips with edge vision and an open-weight model

This is the engineering summary of an open reference architecture. The paper, its object model as JSON and the model register are free to reuse under CC BY 4.0: https://codeatoms.ai/structure-phase-construction-saudi-arabia/. The operator is an illustrative scenario, not a customer.

A precast and site logistics joint venture is in the structure phase of a large mixed use district in the Riyadh region. Fourteen tower cranes, six crawler cranes, a precast yard, a batching plant and a fleet of trucks on one haul road. A slip is found at the Thursday look-ahead meeting, a week after it starts, and a near miss reaches safety on a form filed the next day.

The data to see both earlier already exists. It just lives in twelve places that never meet: Primavera P6, batch plant SCADA, a casting bed spreadsheet, an Excel lift schedule, crane anti-collision logs, haulage GPS, paper gate logs and permits, biometric turnstiles, cameras. Here is how that turns into a system.

1. Join first, model second

Every source enters through one of three adapter families (telemetry, integration, and file and document), never directly, and all three publish onto one Kafka backbone so a paper gate log and a batch ticket arrive with the same guarantees. The adapters write fifteen typed objects: tower crane, crawler crane, casting bed, precast unit, truck, batch ticket, P6 activity, permit, observation form, gate movement, delivery window, lift schedule entry, pour, worker and zone.

The pour is the focal object. Start at a pour marked Delayed and one traversal reaches the batch tickets behind it, the precast units and their curing state, the P6 activity and its float, the delivery windows and trucks, and the cranes that idled while it waited. The same query continues into the zone, its active permits and the workers present. A document store can hold every record and answer none of that.

The object model ships as JSON (ontology/objects.json, format hyper-ontology/1) so you can load it instead of redrawing it.

2. Put each model where its work is cheapest

Tier Hardware Runs Why there
Edge Jetson Orin class in solar powered IP66 enclosures, five locations RF-DETR detection via ONNX Runtime, ByteTrack-class tracking, Frigate NVR, under K3s Private LTE drops at laydown and under crane zones; a breach must be seen where it happens
Site One server, eight L40S-class 48 GB cards Chronos-2 forecasting over InfluxDB, BGE-M3 multilingual retrieval, breach aggregation Forecasting needs site-wide state
Frontier One node, eight 141 GB HBM cards (H200 class) GLM 5.3 at FP8 for the planner work surface Reasoning stays inside the Kingdom beside the record

3. Size the frontier node from the weights

weights = 753B parameters x 1 byte (FP8)   = 753 GB
need    = 753 GB x 1.2 (KV cache, activations) = 904 GB
node    = 8 x 141 GB                        = 1,128 GB  -> 224 GB headroom
Enter fullscreen mode Exit fullscreen mode

Serving the BF16 weights (about 1.5 TB) would need sixteen cards. FP8 halves it to one node.

4. Pick models by licence

Role Model Licence
Reasoning, work surface GLM 5.3 open weights, 753B bespoke; commercial use, fine tuning and redistribution permitted
Detection and segmentation RF-DETR, fine tuned on site footage Apache-2.0, package to checkpoints
Tracking Roboflow trackers (ByteTrack class) Apache-2.0, avoiding AGPL tracker collections
Forecasting Chronos-2 Apache-2.0
Retrieval across a multilingual workforce BGE-M3 MIT

5. Design for heat, dust and a ban on midday work

Dust, glare and 50 °C ambient fail as accuracy problems unless each enclosure carries a thermal budget including solar gain, with lens cleaning as a design parameter. Thermal cameras stay at 9 Hz or below so US-origin units fall outside the export licence for thermal imaging (ECCN 6A003.b.4.b).

6. Keep a person on every write

Every re-sequencing, delivery window and look-ahead draft is a recommendation; the planner or crane coordinator approves it and the approval is recorded with the reasoning. Breach alerts go to the HSE officer, who confirms, dismisses or escalates. The model never closes a gate or stops a lift.

7. What it costs

Three years Cost
Owned, the design above about 642,000 US dollars, including support and power at the Saudi industrial tariff
Rented, the same GPUs around the clock 1.41 million to 2.81 million US dollars
Closed model break-even the cheapest closed model matches the owned stack at about 31 users; above that ownership is cheaper and the gap grows with every user

Ownership lands at about one half the cheapest three-year commitment. Every price in Appendix A is cited with its source and the date it was read.

Full design, figures, Appendix A with every price cited, and the object model: the paper.


Designed on CodeNinja Praxis, CodeNinja's platform for designing physical AI systems. The object model imports into CodeNinja Hyper Ontology, which turns it into a living system. Load it yourself with the open hyper-ontology loader.

Top comments (1)

Collapse
 
aifliproom profile image
AI Flip Room •

Point 5 rings true at a much smaller scale too. With room photos shot on phones, a big share of what first looked like model errors traced back to the input: tilted shots, dark corners, a window blown out by glare. Treating capture conditions as part of the system design, not as a user problem, changed how we read every failure. Is lens cleaning triggered by a drop in detection confidence, or kept on a fixed rota?