DEV Community

Cover image for What 6 Months in Production Taught Me About Biometric Gate Control: Race Conditions, Ghost Fingerprints, and Building a System That Actually Works
Lakshay Tyagi
Lakshay Tyagi

Posted on Originally published at imlakshay08-complete-ruby-on-rails.hashnode.dev

What 6 Months in Production Taught Me About Biometric Gate Control: Race Conditions, Ghost Fingerprints, and Building a System That Actually Works

A follow-up to my previous two posts on connecting a ZK fingerprint device to Rails. This one is about what actually broke in production, how I found it, and what I rebuilt from scratch.

When I published the second biometric post, I genuinely believed the system was solid. Gate control working. Enrollment from the browser. Auto-start on boot. Android backup bridge. The whole thing.

Then the client called.

Expired members were walking in. The dashboard showed DENIED. The door opened anyway. And when I looked at the code — really looked at it — I found the line that explained everything:

from sync_access import sync_device_access  # commentthis
Enter fullscreen mode Exit fullscreen mode

Gate control had never run in production. Not once. Every DENIED label in the dashboard for six months was cosmetic. The door opened for everyone who had a fingerprint on the device, expired subscription or not.

This post is about what I found, what I broke while fixing it, and what the system actually looks like now.


The Mental Model Most People Get Wrong

Before anything else, I need to explain something that took me an embarrassingly long time to fully internalize: the Rails app does not control who enters the gym.

There are two completely separate systems operating independently:

System 1 — The physical gate (device-enforced, zero internet dependency)

The ZK fingerprint device has a door relay connected to it. When someone scans their finger, the device runs a match against the fingerprint templates stored in its own internal flash memory. If it finds a match — green light, relay fires, door opens. If it doesn't — red light, door stays closed.

The device has no knowledge of subscriptions, member names, your database, or the internet. It does not make HTTP requests. It does not check anything in Rails. It simply asks: "do I have a template that matches this finger?" That's the entire decision. Even if Rails is down, Railway is down, the laptop is off, the internet is cut — if a template exists on the device, the door opens.

System 2 — Attendance recording (Rails, happens after the door decision)

After the device fires (or doesn't), the Python bridge picks up the punch and sends it to Rails. Rails looks up the biometric mapping, checks the subscription, and records the attendance as ALLOWED or DENIED. This is purely record-keeping. The door has already opened or stayed closed before this code runs. Rails cannot retroactively close a door.

This distinction matters enormously because it means:

  • trn_member_biometric_mappings does not allow or deny anyone. It's a lookup table. It answers one question: "which gym member does this fingerprint slot belong to?" That's all.
  • trn_member_subscriptions does not control the gate. It tells you who has an active subscription. The gate doesn't know it exists.
  • sync_access.py is the only bridge between the two systems. It translates "subscription expired in Rails DB" into "template deleted from device memory." Without it running, the two systems diverge silently.

The complete lifecycle of one member looks like this:

Enrolled
  → template written to device flash memory
  → DB row created: mbm_is_active='Y', mbm_finger_template=NULL
  → NULL is correct — template is ON device, not backed up yet
  → door opens when they scan ✅

Subscription active
  → sync runs every 60 seconds
  → device_audit returns access=ALLOW
  → sync does nothing
  → template stays on device, NULL stays in DB
  → door opens ✅

Subscription expires
  → sync runs
  → device_audit returns access=DENY
  → sync saves template bytes to DB (mbm_finger_template = JSON)
  → sync deletes user from device (conn.delete_user)
  → template no longer in device memory
  → door stays closed ❌

Member renews
  → sync runs
  → device_audit returns access=ALLOW
  → member not on device but has stored template
  → sync restores template from DB back to device
  → door opens again ✅
Enter fullscreen mode Exit fullscreen mode

mbm_finger_template is NULL for every active member. That's not a bug. It's correct. The template lives on the device. It only moves to the DB right before deletion, so it can be restored on renewal. The DB is a backup, not the primary store.


Gate Control Was Never Running

Back to the commented-out line.

sync_device_access() was disabled in bridge.py during development and never re-enabled before going live. The midnight scheduler thread was starting, running its loop, and doing nothing. For six months.

The first thing I did after re-enabling it was run with --dry-run to see what it would actually do:

Sync complete. Allowed=48, Blocked=0, Orphans=153
Enter fullscreen mode Exit fullscreen mode

Wait. 153 orphans? Those are fingerprint slots on the device with no DB mapping — remnants of manual enrollments, test users, people enrolled before the software existed. And Blocked=0 even though I knew there were expired members?

The dry-run showed nothing to block because of the second bug.


The Sync Blind Spot

Blog 2's sync_access.py looped over mbm_is_active='Y' rows from the DB and checked each one against the device. Logical, but fatally flawed.

When a member re-enrolls, Step 1 of enrollment deletes their old fingerprint slot from the device. Step 2 sets their old mapping to mbm_is_active='N'. Step 3 creates a new slot, Step 4 saves a new active mapping.

The old mapping disappears from the sync's view — it only looks at active rows. But here's the problem: Step 1 runs against the device, but sometimes the old slot had already been deleted in a previous partial enrollment attempt. In those cases, the old slot never got cleaned up properly, or a duplicate slot from an even older enrollment was still sitting on the device.

The result: a member could have multiple fingerprint slots on the device. Sync only knew about the current active one. Old slots — "ghost fingerprints" — were invisible. When a subscription expired, sync blocked the current slot. The ghost slots stayed on the device. The member could still enter with an old finger.

The fix required a completely different approach. Instead of asking the DB "which members exist and are they on the device?", ask the device "which fingerprint slots exist and what's their status?"

New device_audit Rails endpoint returns all mappings — active and inactive — each tagged with is_active_mapping: true/false and access: ALLOW/DENY. The rewritten sync_access.py now walks conn.get_users() (every slot physically on the device), looks each one up in the device_audit results, and acts accordingly:

# PASS 1: walk every fingerprint on device
for user in device_users:
    if user.privilege == 14:
        continue  # never touch staff/admin accounts

    mapping = mappings_by_device_user_id.get(str(user.user_id))

    if not mapping:
        orphan_count += 1
        continue  # no DB record for this slot — leave it alone

    if mapping["access"] == "DENY":
        # save template first, delete second
        save_template_to_rails(...)
        conn.delete_user(uid=user.uid)

# PASS 2: restore anyone who renewed but isn't on device
for mapping in all_mappings:
    if mapping["access"] == "ALLOW" and not on_device and has_template:
        restore_user_to_device(conn, mapping)
Enter fullscreen mode Exit fullscreen mode

After the fix, running the real sync for the first time: 93 expired members blocked in one pass. Ninety-three fingerprints that had been opening the door for months, gone in a single sync cycle.


Two-Phase Save/Delete — The Safety Guarantee

Blog 2's sync had a silent data loss risk:

save_template_to_rails(device_user_id, uid, templates)  # might fail
conn.delete_user(uid=user.uid)  # runs regardless
Enter fullscreen mode Exit fullscreen mode

If the network dropped between these two lines — timeout, cold start on Railway, anything — the template was gone forever. The member's fingerprint was deleted from the device with nothing saved to the DB. No restore possible. Manual re-enrollment required, meaning the member had to physically come in.

The fix is small but critical. save_template_to_rails() now returns True or False:

def save_template_to_rails(device_user_id, uid, templates):
    try:
        response = requests.post(
            f"{RAILS_API_BASE}/api/biometric_mappings/save_template",
            json={...},
            timeout=60
        )
        if response.status_code == 200:
            return response.json().get("status") == True
        return False
    except Exception as e:
        print(f"Template save failed: {e}")
        return False
Enter fullscreen mode Exit fullscreen mode

And the delete only happens if the save was confirmed:

if user_templates and mapping["is_active_mapping"]:
    saved = save_template_to_rails(device_user_id, user.uid, user_templates)
    if not saved:
        print(f"SKIPPED — template save failed, will retry next sync")
        continue  # skip the delete entirely

conn.delete_user(uid=user.uid)
Enter fullscreen mode Exit fullscreen mode

If the save fails, the member is skipped. Their finger stays on the device. The next sync cycle (60 seconds later) tries again. Eventually it succeeds. No data lost.

This is the most important safety property of the system: a subscription expiring cannot cause template loss under any failure condition. Internet down → sync skips, finger stays on device. Bridge down → device untouched. Railway cold start kills the save request → template save returns False, delete skipped, retry next cycle. The only way a template disappears is if the save succeeds AND the delete succeeds.


The Enrollment Race Condition

This section is the one that cost the most debugging hours.

The enrollment flow in Blog 2 computed the next uid like this:

def get_next_uid_and_user_id(conn):
    users = conn.get_users()
    max_uid = max(u.uid for u in users)
    max_user_id = max(int(u.user_id) for u in users)
    return max_uid + 1, str(max_user_id + 1)
Enter fullscreen mode Exit fullscreen mode

Simple. Reasonable. Completely wrong in production.

conn.enroll_user() puts the device into "waiting for scan" mode and returns immediately. The actual fingerprint capture happens whenever the person physically places their finger on the scanner — which could be 5 seconds or 5 minutes later. During that waiting period, conn.get_users() may not show the pending user at all, because no template has been committed yet.

Incident 1 — Concurrent enrollments:

Staff clicked Enroll for member A. Device showed "scan now." Before A scanned, staff clicked Enroll for member B. get_next_uid_and_user_id() called conn.get_users(). A's pending slot wasn't visible. Both A and B computed the same next uid. Both got assigned, say, uid=321. Whoever scanned first "won" the slot. The other person got a mapping pointing to someone else's fingerprint.

In one session, four members ended up sharing uid=321. All four rows in the DB pointing to the same physical slot, one person's actual fingerprint stored there, three others with ghost mappings. Manual cleanup required for all four.

Incident 2 — Process restart between enroll and scan:

Staff clicked Enroll. Before the member scanned, enroll_api.py was restarted to apply a code change. (Critical lesson learned painfully: saving a .py file does NOT update a running Python process. You must stop it and start it again.) The in-memory _last_issued counter reset to zero. The next enrollment computed the same uid as the pending-but-unscanned slot. Collision.

Incident 3 — Old seed data invisible to the DB query:

The max_ids endpoint filtered by mbm_device_sn. Old seed rows from before the system existed had blank or different mbm_device_sn values. They were invisible to the query. max_ids returned 456 as the maximum. An old seed row had device_user_id=457. New enrollment also got assigned 457. Collision.

Three different root causes, same symptom: duplicate uid assignment, ghost mappings, wrong attendance attributed to wrong members.

The fixes, in order of attempt:

  1. Per-member locks — didn't help. Two different members enrolling concurrently still computed the same uid.
  2. Global enrollment lock + _last_issued in-memory counter — fixed concurrent enrollments, but _last_issued reset to zero on every restart.
  3. max_ids endpoint — fixed restart case, but old seed data with wrong device_sn was invisible.
  4. Allocator table with FOR UPDATE row locking — correct final fix.

The allocator table:

CREATE TABLE biometric_id_allocations (
  id INT AUTO_INCREMENT PRIMARY KEY,
  allocation_compcode VARCHAR(12) NOT NULL,
  allocation_device_sn VARCHAR(50) NOT NULL,
  allocation_next_uid INT NOT NULL,
  allocation_next_device_user_id INT NOT NULL,
  created_at DATETIME,
  updated_at DATETIME
);
Enter fullscreen mode Exit fullscreen mode

Every enrollment calls POST /api/biometric_mappings/allocate_ids. Rails runs:

ActiveRecord::Base.transaction do
  row = BiometricIdAllocation
    .lock("FOR UPDATE")
    .find_or_create_by!(
      allocation_compcode: compcode,
      allocation_device_sn: device_sn
    )

  uid = row.allocation_next_uid
  device_user_id = row.allocation_next_device_user_id

  row.update!(
    allocation_next_uid: uid + 1,
    allocation_next_device_user_id: device_user_id + 1
  )
end

render json: { status: true, uid: uid, device_user_id: device_user_id.to_s }
Enter fullscreen mode Exit fullscreen mode

FOR UPDATE locks the row for the duration of the transaction. Two concurrent enrollments cannot both read the same counter value — the second one waits until the first commits, then reads the already-incremented value. Atomic. DB-backed. Survives restarts. Survives concurrent calls. Survives half-enrolled limbo states. No more collisions by construction.

The global enrollment lock is kept alongside it as defense-in-depth — returns HTTP 429 "please wait" if another enrollment is physically in progress — because even with a safe counter, you don't want two people standing at the device simultaneously.

The enrollment function in Python became one line change:

# Before
new_uid, new_device_user_id = get_next_uid_and_user_id(conn)

# After  
new_uid, new_device_user_id = allocate_next_ids(compcode)
Enter fullscreen mode Exit fullscreen mode

Where allocate_next_ids is:

def allocate_next_ids(compcode):
    resp = req.post(
        f"{RAILS_API_BASE}/api/biometric_mappings/allocate_ids",
        json={"compcode": compcode, "device_sn": DEVICE_SN},
        timeout=15
    ).json()
    return resp["uid"], resp["device_user_id"]
Enter fullscreen mode Exit fullscreen mode

Everything else in enrollment — delete old finger, deactivate old mapping, create device user, trigger scan, save new mapping — stayed exactly the same. One line swap. No race conditions.


The Bugs Nobody Talks About

While debugging the race condition, three other silent bugs surfaced.

deactivate_all silently doing nothing:

During re-enrollment, Python sent member_id as an integer in the JSON payload. The DB column mbm_member_id is varchar(4). Rails compared WHERE mbm_member_id = 291 (integer) against '291' (string). Zero rows updated. No error thrown. The old mapping stayed active. The member ended up with two active mappings — the old finger still worked after re-enrollment.

Fix: explicit .to_s.strip in the Rails controller, str(member_id) in Python. One line each.

save_template writing to the wrong row:

When a member had both active and inactive mappings with the same device_user_id (from a failed re-enrollment), find_by returned the first match — often the deactivated old row (lowest id). The template bytes got saved to the wrong row. The active mapping stayed with template=NULL. When the member's subscription expired and sync tried to block them, it found no template to save, skipped the backup, and either left them on the device or deleted them without backup depending on the code path.

Fix: added mbm_is_active: 'Y' to the find_by lookup. Now it only matches the active mapping.

mbm_is_active='N' from birth:

If enrollment fails after Steps 1+2 (old finger deleted, old mapping deactivated) but before Steps 3+4 complete (new device user created, new mapping saved), the system creates a mapping row with mbm_is_active='N' and mbm_finger_template=NULL. The member has a valid subscription but no working fingerprint. The frontend shows "No fingerprint enrolled." Door stays closed.

This is only detectable by querying:

SELECT m.*, s.ms_end_date 
FROM trn_member_biometric_mappings m
JOIN trn_member_subscriptions s ON s.ms_member_id = m.mbm_member_id
WHERE m.mbm_is_active = 'N' 
AND m.mbm_finger_template IS NULL
AND s.ms_end_date >= CURDATE();
Enter fullscreen mode Exit fullscreen mode

Any rows returned are members with active subscriptions and no working fingerprint. Fix: re-enrollment.


Operational Changes

Sync timing — midnight to 60 seconds:

A member who renews their subscription at 2pm shouldn't have to wait until the next morning for the door to work. Sync now runs every 60 seconds. Worst case wait after renewal: 60 seconds. In practice, under a minute.

Device clock sync on startup:

The ZK device has its own internal clock. If it drifts, attendance timestamps are wrong — and the bridge's today-only filter might discard punches that are actually today. Bridge.py now syncs the device clock from the laptop on startup:

def sync_device_time(conn):
    try:
        conn.set_time(datetime.now())
        print(f"Device time synced to {datetime.now().strftime('%Y-%m-%d %H:%M:%S')}")
    except Exception as e:
        print(f"Could not sync device time: {e}")
Enter fullscreen mode Exit fullscreen mode

The laptop syncs automatically with Windows time servers. The device syncs with the laptop every time the bridge restarts. Clock drift eliminated.

Bridge heartbeat and dashboard indicator:

Bridge.py now pings POST /api/bridge_heartbeat every 5 minutes. A new trn_bridge_heartbeats table stores the last-seen timestamp per device. The dashboard shows a dynamic bridge status banner — green if last heartbeat was within 10 minutes, orange if within an hour, red if longer.

The banner auto-updates every 30 seconds via AJAX without page refresh. When the bridge goes offline and staff restart it, they watch the banner flip from red to green on its own.

Offline state shows exactly what to do:

"Biometric Bridge is OFFLINE — Go to gym laptop → double-click 'Restart Biometric Bridge' on Desktop → wait 30 seconds"

restart_bridge.bat on the Desktop:

@echo off
taskkill /F /IM python.exe /T >nul 2>&1
timeout /t 2 /nobreak >nul
start "" /B cmd /c "cd /d C:\Users\Admin\Desktop\biometric_bridge && python bridge.py" >nul 2>&1
start "" /B cmd /c "cd /d C:\Users\Admin\Desktop\biometric_bridge && python enroll_api.py" >nul 2>&1
echo Bridge restarted. Wait 30 seconds then check dashboard.
pause
Enter fullscreen mode Exit fullscreen mode

One double-click. No terminal. No Python knowledge. Staff can handle it.

Device and DB reset scripts:

Two utility scripts written for production operations:

cleanup_selective.py — reads a CSV with a keep column, deletes everyone not marked from the device. Used twice: once when 179 accumulated ghost users needed clearing, once for a full client-requested reset keeping only 5 people (admins and owners).

deactivate_orphaned_mappings.py — after a device wipe, queries device_audit and deactivates DB rows pointing to slots that no longer exist on device. Keeps DB and device in sync after bulk operations.


Why Enterprise Systems Don't Have These Problems

Brief and honest: eSSL, ZKTeco's own software, Gymmaster, and similar systems avoid most of these issues because they use the official ZKTeco Windows SDK (zkemkeeper.dll), not pyzk.

The official SDK has enrollment event callbacks — when a scan completes, an event fires in your application with the template bytes. You save atomically in one step. No fire-and-forget, no polling, no timing window for race conditions.

They also use the ADMS protocol — the device itself initiates an HTTP connection to the cloud and pushes attendance records. No bridge laptop. No polling. I investigated ADMS and got close — DNS worked, port 80 was reachable, Rails controller was deployed — but the router at the gym appears to have per-device firewall rules that block outbound HTTP from the ZK device's MAC address specifically. The laptop and Android phone on the same network work fine. The device can't reach the internet. ADMS code stays in the codebase waiting for a router upgrade.

The enrollment race condition exists because pyzk has no enrollment callback — a fundamental limitation of a community-reverse-engineered library versus the official SDK. The allocator table is the correct workaround within those constraints. It works.


Current State

Enrollment       → allocator table, FOR UPDATE locking, zero collisions
Sync             → every 60 seconds, walks device not DB, 
                   catches all fingerprints including ghosts,
                   two-phase save/delete, retry on failure
Gate control     → device-enforced, internet-independent,
                   template on device = door opens, always
Template safety  → saved before delete, confirmed before delete,
                   retry if save fails, no irreversible loss possible
Monitoring       → heartbeat every 5 min, dashboard indicator,
                   non-technical restart procedure on Desktop
Device clock     → synced to laptop time on every bridge startup
Enter fullscreen mode Exit fullscreen mode

The first sync run after re-enabling gate control blocked 93 fingerprints. Two months later the system has processed thousands of attendance records, correctly blocked expired members, correctly restored renewed ones, and handled two full device resets without data loss.

The notebook with the manual attendance register is still in the drawer. It has not been opened.


Key Takeaways

The hardest bugs aren't the ones that throw errors. They're the ones that silently do nothing while you think everything is working. sync_device_access() commented out. deactivate_all returning updated: 0 with no exception. Template save failing silently and delete running anyway. None of these crashed the system. All of them were invisible until something downstream went wrong.

"Works" and "production-ready" are different things. The system worked in testing. It worked for the first few weeks. The bugs only appear at scale, over time, with real staff who click buttons too fast, real network conditions that drop connections mid-request, and real device firmware quirks that don't show up in controlled environments.

Log everything that matters. deactivate_all called: compcode=SF member_id=291 and deactivate_all: updated 1 rows — two log lines that took 20 minutes to add and saved hours of debugging. If a critical operation can return 0 when it should return 1, log it explicitly. Don't assume.

The device is the source of truth for physical access. Not Rails. Not the DB. The device. Everything else is record-keeping. Get this mental model right and the architecture becomes obvious.


Live: spine-fitness.com | Source: github.com/imlakshay08/spine-fitness-gym-management-system

Part of a series — read the original post and the first upgrade post for full context.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to