Hi DEV community,
We are currently refining the edge communication layer for an enterprise-grade, multi-facility access control system deployment and running into an architectural debate regarding edge controller resilience during WAN network partitions.
The Setup & Architecture
Edge Nodes: Embedded Linux controllers interfacing with biometric terminals and encrypted RFID card readers (MIFARE DESFire EV3) over OSDP v2.2 (RS-485).
Physical Actuation: 12V fail-secure door strike relays driven via optocoupled GPIO pins.
Requirements: Sub-200ms badge-to-unlock latency, with zero downtime if the central cloud backend drops offline.
The Challenge
When internet connectivity drops, edge readers must authenticate credentials locally using an on-device credential cache. However, we're encountering two specific edge cases:
Cache Invalidation & Real-time Revocation: If a credential or badge is revoked centrally while a remote branch loses WAN connectivity, pushing delta updates via MQTT stalls. How do you mitigate the security window between WAN restoration and cache reconciliation?
Audit Buffer Flooding: High-throughput turnstiles and gates generate bursts of badge read events during morning peak hours (8:00 AM – 9:00 AM). When offline, storing thousands of encrypted transit events and re-streaming them post-reconnect causes brief CPU spikes that delay subsequent real-time badge validations.
For those who have architected mission-critical edge IoT or security systems:
Do you separate the real-time relay actuation daemon onto a dedicated RTOS/microcontroller while running MQTT/sync on a separate application core?
What lightweight embedded storage engine (SQLite, RocksDB, LMDB) do you prefer for high-concurrency read/write operations without flash storage degradation?
Would love to hear your architectural patterns and lessons learned!
Top comments (0)