DEV Community

Cover image for Building a License Authentication System From Scratch
MR storm
MR storm

Posted on

Building a License Authentication System From Scratch

Or: how to defend a desktop app when you've already handed the attacker the source code.


Table of contents

  1. The problem
  2. Why boolean checks fail
  3. The core idea: server-gated data
  4. The full architecture
  5. The wire protocol
  6. Picking your stack
  7. Step 1 — Init and the encrypted payload
  8. Step 2 — Password storage done right
  9. Step 3 — Register and license activation
  10. Step 4 — Login and session creation
  11. Step 5 — Hardware binding
  12. Step 6 — Heartbeat and rotating tokens
  13. Step 7 — Fetching the gated data
  14. Step 8 — Anti-relay with sequence counters
  15. Step 9 — Binary integrity
  16. Step 10 — SDK version gating
  17. Step 11 — Rate limiting
  18. Step 12 — Failed login tracking and auto-blacklist
  19. Step 13 — Login anomaly detection
  20. Step 14 — The admin side
  21. Step 15 — The client side
  22. What's real defense and what's theater
  23. Common mistakes
  24. Testing your system
  25. The honest ending

1. The problem

If you've ever shipped a desktop application that people pay for, you've run into this: the moment your code runs on someone else's machine, you've lost control of it.

You can compile it, pack it, obfuscate it, but at the end of the day there's a binary on someone's disk, and that binary is theirs. They can read it, they can patch it, they can run it in a debugger, they can feed it fake network responses, they can do anything they want with it. And every check you put inside that binary to enforce payment runs inside code they control.

This is the fundamental asymmetry of software licensing. The defender has to get every check right. The attacker only has to find one hole.

The good news is that this asymmetry doesn't mean you're helpless. It means you have to shift your thinking. Instead of trying to protect the decision (is this user licensed?), you protect the data (what does a licensed user get that an unlicensed user doesn't have?).

That shift — from protecting decisions to protecting data — is what this entire post is about.


2. Why boolean checks fail

Start with the obvious approach. Somewhere in your code, you write:

def is_licensed():
    # check something
    return True

if not is_licensed():
    exit(1)

# premium features below
Enter fullscreen mode Exit fullscreen mode

Anyone with a hex editor and thirty seconds finds that if not is_licensed(): branch and replaces it with a no-op. Now the app runs whether or not it's licensed. Your entire protection was one byte.

You can make it harder to find. Obfuscate the function name, scatter the check across multiple files, compute it from multiple values. All of that helps a little. All of it fails eventually.

Why? Because the result of the check has to be used by the client. And the client is attacker-controlled. The moment the app decides to enable a feature based on a value that lives inside the attacker's memory, the attacker can change that value. It doesn't matter how cleverly you computed it. It only matters that it's local.

This is the first law of client-side security: any decision the client makes, the client can lie about.

You can't get around this by being clever. You have to change the game.


3. The core idea: server-gated data

Here's the shift.

Instead of your app asking "am I licensed?" and getting a yes/no, your app asks the server for the data it actually needs to function. If the server says no, the app doesn't get a boolean False it can patch to True. It gets nothing. And nothing is unfakeable.

Concretely. Imagine a tool with a premium feature that generates reports. The naive design:

if user.is_premium:
    template = load_local_template()
    generate_report(template)
Enter fullscreen mode Exit fullscreen mode

The secure design:

encrypted = server.fetch_gated_data()
template = decrypt(encrypted, local_key)
generate_report(template)
Enter fullscreen mode Exit fullscreen mode

The attacker patches if user.is_premium — nothing happens. That was never the gate. The gate is that template doesn't exist on their machine in usable form until the server hands it over, encrypted, tied to their session, bound to their hardware.

They can't patch their way into data they don't have. They can only:

  • Steal a legit user's session and use it live (which you can detect)
  • Break the encryption (which is infeasible with proper crypto)
  • Run a proxy in front of a legit client (which you can defeat with sequence counters)
  • Get the actual data some other way (which is a social problem, not a technical one)

This is the single most important idea in modern license systems. Every technique below is defense in depth around this idea. The idea itself is the foundation.

If your system doesn't have server-gated data, nothing else in this post matters. You're polishing a broken design.


4. The full architecture

Before we write code, look at the whole picture.

┌──────────────┐          ┌──────────────┐          ┌──────────────┐
│   CLIENT     │          │   SERVER     │          │   DATABASE   │
│  (desktop)   │          │  (FastAPI)   │          │  (SQL/JSON)  │
└──────┬───────┘          └──────┬───────┘          └──────┬───────┘
       │                         │                         │
       │ 1. init                 │                         │
       │────────────────────────▶│                         │
       │                         │ create session          │
       │                         │ generate enc_key        │
       │                         │ store key server-side   │
       │◀────────────────────────│                         │
       │ sessionid, enc_payload  │                         │
       │                         │                         │
       │ 2. login (user,pass,hwid)                         │
       │────────────────────────▶│                         │
       │                         │ verify password ────────▶
       │                         │◀────────────────────────│
       │                         │ check HWID              │
       │                         │ bind session to hwid    │
       │◀────────────────────────│                         │
       │ decryption_key, session │                         │
       │                         │                         │
       │ 3. heartbeat            │                         │
       │────────────────────────▶│                         │
       │                         │ rotate hb_token         │
       │◀────────────────────────│                         │
       │ hb_token                │                         │
       │                         │                         │
       │ 4. fetchdata (hb_token, fetch_seq)                │
       │────────────────────────▶│                         │
       │                         │ verify token            │
       │                         │ verify seq > last_seq   │
       │                         │ encrypt with hwid_key   │
       │◀────────────────────────│                         │
       │ encrypted_data          │                         │
       │                         │                         │
       │ decrypt with derived key│                         │
       │ use data                │                         │
Enter fullscreen mode Exit fullscreen mode

The client never receives the raw gated data. It receives encrypted data that only its own hardware can decrypt. The server never trusts the client's claims about itself — every request re-verifies HWID, session, token, and sequence.


5. The wire protocol

Before writing a line of server code, decide what the client and server actually say to each other. This is your contract, and changing it later is painful because every deployed client breaks.

Every request from the client looks roughly like:

{
  "type": "login",
  "name": "my-app-name",
  "ownerid": "abc123def456",
  "username": "user1",
  "password": "hunter2",
  "hwid": "9f3a2c7e1b4d...",
  "sessionid": "optional, set after init",
  "hb_token": "optional, set after first heartbeat",
  "fetch_seq": 0
}
Enter fullscreen mode Exit fullscreen mode

And every response is JSON with at minimum:

{
  "success": true,
  "message": "human-readable",
  ...action-specific fields
}
Enter fullscreen mode Exit fullscreen mode

Design the protocol so that every request carries the full context needed to verify it. Don't have a "stateful" protocol where the client says one thing on request 1 and another on request 2 — you'll get desync bugs. Every request should be verifiable on its own.

A minimal set of action types:

Action Purpose
init Start a session, get an encrypted payload the client can't yet decrypt
register Create a user with a license key
login Authenticate an existing user with username/password
license Authenticate with just a key (no username/password)
heartbeat Prove the session is alive, get a new rotating token
fetchdata Retrieve the gated payload, encrypted per-HWID

Optional, add later:

Action Purpose
check Lightweight session validity check
setvar / getvar Per-user variables stored server-side
log Client-side event logging
webhook Trigger outbound notifications

That's it. Everything else is a variation. Keep the protocol small.


6. Picking your stack

Server: FastAPI over Flask. Both work. FastAPI gives you:

  • Async request handling out of the box
  • Typed request models via Pydantic (so the request schema is the documentation)
  • Automatic validation (malformed requests get rejected before your code runs)
  • Better performance under concurrent load

For an auth server that gets hit constantly by heartbeats from every client, that last point matters. Flask with sync workers will run out of worker threads at low client counts. FastAPI handles thousands of concurrent requests per worker.

Storage: a real database, or careful file locking. JSON files are fine for a prototype. They stop being fine the moment two users register at the same time, because you get a race condition:

User A: read keys.json  → key "ABC123" is unused
User B: read keys.json  → key "ABC123" is unused
User A: mark ABC123 used → write keys.json
User B: mark ABC123 used → write keys.json  (overwrites A's write)
Enter fullscreen mode Exit fullscreen mode

Both users think they registered with the same key. The second write wins, and now the audit log is wrong.

Two fixes:

  1. Use SQLite. Transactions handle this for you. Add Postgres when you outgrow it.
  2. If you must use JSON files, wrap every read-modify-write in a per-resource lock:
import threading
app_locks = {}
locks_guard = threading.Lock()

def get_app_lock(app_id):
    lock = app_locks.get(app_id)
    if lock is None:
        with locks_guard:
            lock = app_locks.setdefault(app_id, threading.Lock())
    return lock

# Every mutation:
with get_app_lock(app_id):
    data = read_json(f"app_{app_id}_keys.json")
    # ... mutate ...
    write_json(f"app_{app_id}_keys.json", data)
Enter fullscreen mode Exit fullscreen mode

The lock only works within one process. If you run multiple server workers, they don't share locks and you're back to races. Either pin to one worker or use a real database.

Client: whatever the app is written in. The protocol is HTTP + JSON. The client just has to implement it. The decryption step requires HMAC and hashing primitives, which every language has.


7. Step 1 — Init and the encrypted payload

The very first thing a client does is init. The server:

  1. Generates a random session ID
  2. Generates a random encryption key for this session only
  3. Encrypts a small metadata payload with that key
  4. Stores the key server-side, keyed by the session ID
  5. Returns the session ID and the encrypted blob, but not the key
import secrets, base64, hmac, hashlib, json, time

_enc_sessions = {}  # session_id -> {enc_key, mac_key, app_id, created, ...}

def encrypt_payload(session_id, app_id, payload):
    enc_key = secrets.token_bytes(32)
    mac_key = secrets.token_bytes(32)
    iv = secrets.token_bytes(16)

    plaintext = json.dumps(payload, separators=(",", ":")).encode()
    # PKCS7 pad to block boundary (even though we're using a stream, this
    # hides the plaintext length from anyone analyzing ciphertext size)
    pad_len = 16 - (len(plaintext) % 16)
    plaintext += bytes([pad_len]) * pad_len

    ciphertext = hmac_stream(enc_key, iv, plaintext)
    mac = hmac.new(mac_key, iv + ciphertext, hashlib.sha256).digest()

    _enc_sessions[session_id] = {
        "enc_key": enc_key,
        "mac_key": mac_key,
        "app_id": app_id,
        "created": time.time(),
        "authenticated": False,
        "hwid": None,
        "hb_token": None,
        "last_hb": None,
        "client_seq": 0,
    }
    return base64.b64encode(iv + ciphertext + mac).decode()


def hmac_stream(key, iv, data):
    """HMAC-SHA256 counter-mode keystream, XORed with the data."""
    out = bytearray()
    for i in range((len(data) + 31) // 32):
        block = hmac.new(key, iv + i.to_bytes(4, "big"), hashlib.sha256).digest()
        out.extend(block)
    return bytes(a ^ b for a, b in zip(data, out[:len(data)]))
Enter fullscreen mode Exit fullscreen mode

A note on the cipher, because this will get flagged: rolling your own stream cipher is generally a bad idea. The construction above — HMAC-SHA256 in counter mode — is a reasonable PRF-based keystream, but it's not what cryptography engineers would reach for. If you're using the cryptography library for anything else (like Fernet, which is what most people use for at-rest encryption), you already have access to AES-GCM. Use it. It's authenticated, it's constant-time in the important places, it's been reviewed by everyone.

The reason I'm showing the HMAC construction is because it's what a lot of hobbyist auth systems end up writing, and it's instructive to see why it exists (no dep, simple to audit) and why it's inferior to AES-GCM (no formal security proof, easy to misuse, no nonce-misuse resistance).

If you want the honest short version: use AES-GCM. Everything below assumes that level of security.

The init handler itself:

def handle_init(name, ownerid, ver):
    app = find_app_by_name_and_owner(name, ownerid)
    if not app:
        return {"success": False, "message": "Application not found"}
    if not app.get("enabled", True):
        return {"success": False, "message": "Application is disabled"}

    session_id = secrets.token_hex(16)
    payload_data = {
        "app_id": app["id"],
        "app_name": app["name"],
        "version": app.get("version", "1.0.0"),
        "session_id": session_id,
        "nonce": secrets.token_hex(8),
        "ts": datetime.utcnow().isoformat() + "Z",
    }
    encrypted = encrypt_payload(session_id, app["id"], payload_data)

    return {
        "success": True,
        "message": "Initialized",
        "sessionid": session_id,
        "encrypted_payload": encrypted,
        "appinfo": {
            "version": app.get("version", "1.0.0"),
            "numUsers": count_users(app["id"]),
            "numKeys": count_keys(app["id"]),
        },
    }
Enter fullscreen mode Exit fullscreen mode

The client now has a blob it can't read. It needs to authenticate to get the key. The init step is mostly theater on its own — the payload is just metadata — but it establishes the pattern: the client can't proceed without server cooperation.

Also notice: the server generates a fresh random key per session. This is important. It means if an attacker compromises one session key, they've compromised one session, not the entire app. The key never leaves the server until auth succeeds, and even then it's returned as a value that only works for this session.


8. Step 2 — Password storage done right

This is where almost every tutorial gets it wrong, and where I've seen more real-world breaches than anywhere else. So let's be blunt about it.

You do not encrypt passwords. You don't wrap them in symbols. You don't encrypt-then-hash.

Encryption is reversible by design. If you encrypt a password to store it, you must be able to decrypt it later — which means the key lives somewhere on your server. Any attacker who gets your database and your server code gets every password in plaintext. You've gained nothing.

You hash them. Specifically, you use a slow hash with a per-user salt.

Why hashing and not encryption

Hashing is one-way. You store hash(password), not password. When a user logs in, you compute hash(attempted_password) and compare it to the stored hash. If they match, the password was right.

If an attacker steals your database, they get hashes. Not passwords. And if you did it right, they can't reverse the hashes without brute-forcing each one, which is what we're going to make prohibitively expensive.

Why a salt

Without a salt, the hash of "hunter2" is the same for every user. An attacker can precompute a table:

"password"     → 5f4dcc3b5aa765d61d8327deb882cf99
"hunter2"      → 2ab96390c7dbe3439de74d0c9b0b1767
"letmein"      → 0d107d09f5bbe40cade3de5c71e9e9b7
...10 billion more...
Enter fullscreen mode Exit fullscreen mode

Then every hash in your stolen database is a lookup. That's a rainbow table attack.

A salt is a random value generated per user, stored alongside the hash. You hash salt + password. Now every user's "hunter2" produces a different hash. The attacker has to brute-force each user individually, and precomputation is worthless.

Why a slow hash

SHA-256 was designed for speed. A modern GPU can compute billions of SHA-256 hashes per second. If your password has 8 characters from a 62-character alphabet, that's ~218 trillion combinations. Sounds like a lot, but at 10 billion hashes/second, that's 6 hours.

bcrypt, argon2, and scrypt are deliberately slow. They have a "cost" parameter that scales their runtime. Set bcrypt to cost 12 and each hash takes ~250ms. That turns 6 hours into 1.7 million years. Even scaling up the GPU farm doesn't help — the algorithms are memory-hard and don't parallelize the way SHA-256 does.

The right way

import bcrypt

def hash_password(password: str) -> str:
    return bcrypt.hashpw(
        password.encode("utf-8"),
        bcrypt.gensalt(rounds=12),
    ).decode("utf-8")


def verify_password(password: str, stored: str) -> bool:
    try:
        return bcrypt.checkpw(password.encode("utf-8"), stored.encode("utf-8"))
    except Exception:
        return False
Enter fullscreen mode Exit fullscreen mode

That's it. bcrypt.gensalt handles salt generation, bcrypt.hashpw handles the hashing, bcrypt.checkpw handles verification. No manual salt, no manual comparison, no timing leaks.

bcrypt hashes have a known format: they start with $2a$, $2b$, or $2y$ followed by the cost, salt, and hash all in one string. So you can detect bcrypt hashes just by looking at them:

def is_bcrypt(stored):
    return isinstance(stored, str) and stored.startswith(("$2a$", "$2b$", "$2y$"))
Enter fullscreen mode Exit fullscreen mode

Migrating from a worse scheme

If you already have users stored with SHA-256 or plaintext (maybe from an early prototype), you can't easily migrate them all at once — you don't have the plaintext passwords, so you can't re-hash them.

The trick is lazy migration: on each successful login, re-hash with the new scheme and save.

def verify_and_upgrade(password, user):
    stored = user["password_hash"]
    if is_bcrypt(stored):
        return bcrypt.checkpw(password.encode(), stored.encode()), False

    if is_sha256_hex(stored):
        ok = hmac.compare_digest(stored, hashlib.sha256(password.encode()).hexdigest())
        return ok, ok  # needs rehash if matched

    # Legacy plaintext
    ok = hmac.compare_digest(stored, password)
    return ok, ok

# In the login handler:
ok, needs_rehash = verify_and_upgrade(req.password, user)
if ok and needs_rehash:
    user["password_hash"] = hash_password(req.password)
    save_user(user)
Enter fullscreen mode Exit fullscreen mode

Over time every active user migrates. Inactive users get migrated on their next login. Everyone else is old and irrelevant.


9. Step 3 — Register and license activation

Registration creates a user account and burns a license key. Two writes have to happen atomically:

  1. Append the new user to the user list
  2. Mark the license key as used

If either fails or if another registration races in between, you've either created a user without burning a key (free account), or burned a key without creating a user (customer loses their key). Both are bad.

Serialize the whole thing under a per-app lock:

def handle_register(req, ip):
    app_data, app_id = resolve_app(req)
    if not app_data:
        return {"success": False, "message": "Invalid app"}

    with get_app_lock(app_id):
        users = get_app_data(app_id, "users", [])
        keys = get_app_data(app_id, "keys", [])
        blacklist = get_app_data(app_id, "blacklist", [])

        # Check blacklists first (IP and HWID)
        if is_blacklisted(blacklist, "ip", ip):
            return {"success": False, "message": "IP is blacklisted"}
        if is_blacklisted(blacklist, "hwid", req.hwid):
            return {"success": False, "message": "HWID is blacklisted"}

        # Username collision
        if any(u["username"] == req.username for u in users):
            return {"success": False, "message": "Username already taken"}

        # Constant-time key lookup. hmac.compare_digest prevents timing
        # oracles — an attacker can't tell how "close" their guess was.
        req_key = req.key or ""
        key_idx = next(
            (i for i, k in enumerate(keys)
             if hmac.compare_digest(k.get("key", ""), req_key)),
            None,
        )
        if key_idx is None:
            return {"success": False, "message": "Invalid license key"}

        key_data = keys[key_idx]
        if key_data.get("used_by"):
            return {"success": False, "message": "License key already used"}

        # Compute expiry based on key's duration
        duration = key_data.get("duration", "1 month")
        td = DURATION_MAP.get(duration)
        sub_expiry = (datetime.utcnow() + td).isoformat() + "Z" if td else "lifetime"

        new_user = {
            "username": req.username,
            "password_hash": hash_password(req.password),
            "hwid": req.hwid,
            "ip": ip,
            "created_at": datetime.utcnow().isoformat() + "Z",
            "last_login": datetime.utcnow().isoformat() + "Z",
            "subscription": duration,
            "sub_expiry": sub_expiry,
            "banned": False,
            "session_id": secrets.token_hex(16),
            "user_vars": {},
            "total_logins": 1,
        }
        users.append(new_user)
        keys[key_idx]["used_by"] = req.username
        keys[key_idx]["used_at"] = datetime.utcnow().isoformat() + "Z"

        save_app_data(app_id, "users", users)
        save_app_data(app_id, "keys", keys)

    # Outside the lock: slow I/O that doesn't touch app data
    log_event(app_id, "auth", "info", f"User registered: {req.username}", req.username, ip)
    notify_customer(app_id, f"New user: {req.username}")
    return {"success": True, "message": "Registration successful", "info": user_info(new_user)}
Enter fullscreen mode Exit fullscreen mode

A few things to notice:

The lock is per-app, not global. Concurrent registrations on different apps don't block each other. This matters when you have many apps and high traffic.

The lock covers only the mutation. Logging, Telegram notifications, and other I/O happen outside the lock. If you hold a lock while making a 10-second network call to Telegram, every other registration on that app stalls for 10 seconds. That's how you turn one slow API into a DDoS on yourself.

The key comparison uses hmac.compare_digest. This is a constant-time comparison. A naive k["key"] == req_key leaks information via timing — an attacker can learn the correct key byte by byte by measuring how long the comparison takes. compare_digest always takes the same time regardless of how much of the string matches.

The HWID is required. If req.hwid is empty, _validate_hwid rejects before we even get here. This is critical. Early in my own build I had a bug where an empty HWID skipped the HWID check entirely — meaning anyone could bypass hardware binding by omitting the field.


10. Step 4 — Login and session creation

Login is the most attacked endpoint in any auth system. Credential stuffing, brute forcing, and key enumeration all hit it. The handler has to do a lot.

def handle_login(req, ip):
    app_data, app_id = resolve_app(req)
    if not app_data:
        return {"success": False, "message": "Invalid app"}

    # Deferred events — set inside the lock, fired after release
    auto_blacklisted = None
    login_success = False
    user = None
    killed_sessions = 0

    with get_app_lock(app_id):
        users = get_app_data(app_id, "users", [])
        blacklist = get_app_data(app_id, "blacklist", [])

        if is_blacklisted(blacklist, "ip", ip):
            return {"success": False, "message": "IP is blacklisted"}
        if is_blacklisted(blacklist, "hwid", req.hwid):
            return {"success": False, "message": "HWID is blacklisted"}

        # Find user
        user_idx = next(
            (i for i, u in enumerate(users) if u["username"] == req.username),
            None,
        )
        if user_idx is None:
            # Don't reveal whether the username exists or the password
            # was wrong. Same message for both.
            _record_failed_login(app_id, ip, req.hwid)
            return {"success": False, "message": "Invalid credentials"}

        user = users[user_idx]
        ok, needs_rehash = verify_password(req.password, user.get("password_hash", ""))

        if not ok:
            _record_failed_login(app_id, ip, req.hwid)
            auto_blacklisted = _maybe_auto_blacklist(app_id, ip)
            failed_resp = {"success": False, "message": "Invalid credentials"}
        else:
            # Clear the failure counter on success
            _failed_logins.pop((app_id, ip), None)

            if user.get("banned"):
                return {"success": False, "message": f"Banned: {user.get('ban_reason', 'N/A')}"}

            if user["sub_expiry"] != "lifetime":
                try:
                    exp = datetime.fromisoformat(user["sub_expiry"].replace("Z", ""))
                    if datetime.utcnow() > exp:
                        return {"success": False, "message": "Subscription expired"}
                except Exception:
                    pass

            # HWID check + migration
            max_hwids = app_data.get("max_hwids", 1)
            stored_hwid = user.get("hwid", "")
            if req.hwid:
                if stored_hwid and stored_hwid != req.hwid and max_hwids <= 1:
                    return {"success": False, "message": "HWID mismatch. Contact support."}
                if not stored_hwid:
                    user["hwid"] = req.hwid

            # Rotate session ID — the old one dies now
            old_session = user.get("session_id") or ""
            user["session_id"] = secrets.token_hex(16)
            user["last_login"] = datetime.utcnow().isoformat() + "Z"
            user["ip"] = ip
            user["online"] = True
            user["total_logins"] = user.get("total_logins", 0) + 1

            if needs_rehash:
                user["password_hash"] = hash_password(req.password)

            users[user_idx] = user
            save_app_data(app_id, "users", users)

            # Kill any encryption sessions still bound to the old session ID
            if app_data.get("kill_prior_sessions", True) and old_session:
                killed_sessions = kill_prior_sessions(old_session)

            login_success = True

    # Outside the lock
    if not login_success:
        log_event(app_id, "auth", "warning", f"Failed login: {req.username}", req.username, ip)
        if auto_blacklisted:
            log_event(app_id, "security", "error",
                      f"Auto-blacklisted {ip} after repeated failures", req.username, ip)
        return failed_resp

    log_event(app_id, "auth", "info", f"User logged in: {req.username}", req.username, ip)
    if killed_sessions:
        notify_customer(app_id, f"Killed {killed_sessions} prior sessions for {req.username}")

    # Generate the session decryption key
    dec_key = None
    if req.sessionid and req.sessionid in _enc_sessions:
        dec_key = make_decryption_key(req.sessionid, req.hwid)
        _enc_sessions[req.sessionid]["user_session"] = user["session_id"]

    return {
        "success": True,
        "message": "Login successful",
        "info": user_info(user),
        "sessionid": user["session_id"],
        "decryption_key": dec_key,
        "heartbeat_interval": HEARTBEAT_INTERVAL,
    }
Enter fullscreen mode Exit fullscreen mode

A few things worth pointing out:

"Invalid credentials" for both unknown users and wrong passwords. If you return "username not found" for one and "wrong password" for the other, you've built a username enumerator. Attacker fires 10,000 usernames and learns which ones exist. Always return the same message.

Session rotation on every login. When a user logs in, the previous session ID is invalidated and a new one is created. If an attacker stole the old session ID, it dies the moment the legit user logs in again. This is a strong defense against session sharing.

Killing prior encryption sessions. After rotating the user's session ID, any _enc_sessions entries still pointing to the old ID are evicted. A relay holding the old session can't keep fetching data — its session is gone.

The order of operations matters. Check blacklist → find user → verify password → check banned → check subscription → check HWID → rotate session → save. Never commit partial state. If any check fails, return without having written anything.


11. Step 5 — Hardware binding

Every authenticated request carries a hardware ID — a fingerprint derived from the machine's physical properties. The user's session becomes bound to that HWID. Any subsequent request from a different HWID is rejected.

Generating the HWID

The client computes something like:

import platform, uuid, hashlib, hmac

def get_hwid():
    parts = [
        platform.node(),      # hostname
        platform.machine(),   # CPU architecture
        str(uuid.getnode()),  # MAC address (best-effort)
        platform.processor(),
    ]
    raw = "|".join(parts).encode()
    # HMAC so it's not a plain hash of well-known inputs
    return hmac.new(b"hwid-salt", raw, hashlib.sha256).hexdigest()
Enter fullscreen mode Exit fullscreen mode

On Linux you might also read /etc/machine-id. On Windows, wmic csproduct get uuid. The exact inputs don't matter much — what matters is stability. If the HWID changes when the user updates their GPU driver, you'll get false positives.

Server-side checks

Three rules for HWID comparison:

1. Use constant-time comparison.

if not hmac.compare_digest(stored_hwid, incoming_hwid):
    return reject()
Enter fullscreen mode Exit fullscreen mode

Not ==. The difference is that == on strings short-circuits as soon as a byte differs, so the response time leaks how many bytes matched. An attacker can learn a 64-byte HWID one byte at a time with about 64 × 256 requests. compare_digest always takes the same time.

2. Reject empty HWIDs explicitly.

if not stored_hwid or not incoming_hwid:
    return reject()
Enter fullscreen mode Exit fullscreen mode

The bug to avoid: if stored_hwid and not matches(stored, incoming): reject(). If incoming is empty and stored is set, the outer condition is stored_hwid and not matches(...). matches would need to return True for an empty incoming, which most implementations do (because "" == ""). So the check passes. Free bypass.

Make the rule: empty on either side = reject. No exceptions.

3. Support format migration.

If you ever change the HWID format (e.g. from 16 hex chars to 64), you need to accept the old format for existing users while enforcing the new format for new ones. Otherwise every existing user is locked out.

def hwid_matches(stored, incoming):
    if not stored or not incoming:
        return False
    if hmac.compare_digest(stored, incoming):
        return True
    # Migration window: old (16-char) stored, new (64-char) incoming
    if len(stored) == 16 and len(incoming) == 64:
        return True  # accept, and upgrade stored value
    return False
Enter fullscreen mode Exit fullscreen mode

Then when the check passes via migration, update user["hwid"] = incoming. The next login uses the new format.


12. Step 6 — Heartbeat and rotating tokens

Once authenticated, the client starts a timer. Every N seconds (typically 20-60), it sends a heartbeat. The server:

  1. Confirms the session still exists
  2. Confirms the HWID still matches
  3. Confirms the subscription hasn't expired mid-session
  4. Generates a new random hb_token, stores it on the session, and returns it
  5. Updates last_hb timestamp
HEARTBEAT_INTERVAL = 20  # seconds

def validate_heartbeat(session_id, hwid):
    session = _enc_sessions.get(session_id)
    if not session:
        return {"success": False, "message": "Invalid session"}

    if not session["authenticated"]:
        return {"success": False, "message": "Not authenticated"}

    # HWID must match — empty counts as mismatch
    if session["hwid"] and not hwid_matches(session["hwid"], hwid):
        _enc_sessions.pop(session_id, None)
        return {"success": False, "message": "HWID mismatch"}

    # Absolute session lifetime
    if time.time() - session["created"] > SESSION_MAX_AGE:
        _enc_sessions.pop(session_id, None)
        return {"success": False, "message": "Session expired"}

    # Cross-check against the database: is the user still valid?
    app_id = session["app_id"]
    user_session = session.get("user_session")
    if user_session:
        users = get_app_data(app_id, "users", [])
        user = next((u for u in users if u.get("session_id") == user_session), None)

        if not user:
            _enc_sessions.pop(session_id, None)
            return {"success": False, "message": "User no longer exists"}

        if user.get("banned"):
            _enc_sessions.pop(session_id, None)
            return {"success": False, "message": "User banned"}

        if user["sub_expiry"] != "lifetime":
            try:
                exp = datetime.fromisoformat(user["sub_expiry"].replace("Z", ""))
                if datetime.utcnow() > exp:
                    _enc_sessions.pop(session_id, None)
                    return {"success": False, "message": "Subscription expired"}
            except Exception:
                pass

    # Rotate the heartbeat token
    new_token = secrets.token_hex(16)
    session["hb_token"] = new_token
    session["last_hb"] = time.time()
    session["hb_count"] = session.get("hb_count", 0) + 1

    return {
        "success": True,
        "message": "Heartbeat OK",
        "hb_token": new_token,
        "next_heartbeat": HEARTBEAT_INTERVAL,
    }
Enter fullscreen mode Exit fullscreen mode

Why rotate the token?

The rotating token is what makes replay attacks impractical. Suppose an attacker captures a valid fetchdata request. If the token is static, they can replay that request forever within the session window. With rotation:

  • The legit client sends fetchdata with token T1
  • The attacker replays the same request 5 seconds later
  • In between, the legit client sent a heartbeat, and the server rotated the token to T2
  • The attacker's replay now presents T1, which is stale
  • Server rejects

The window of usefulness for a captured request is one heartbeat interval — typically 20-60 seconds. Long enough for legitimate network jitter, short enough that replay is a dead end.

The client's obligation

Every request after the first heartbeat must include the current hb_token. The client's job is to:

  1. Parse hb_token from the heartbeat response
  2. Store it
  3. Include it in every subsequent fetchdata (and any other authenticated call)

If the client fails to include it, the server rejects with a message that tells them to update their SDK.

What this defends against

  • Session replay — captured requests go stale in seconds
  • Session sharing — two clients sharing a session can't both keep up with token rotation
  • Delayed attacks — even if an attacker captures a full session, they can't reuse it after the legitimate client moves on

What this does not defend against

A relay. A relay is a proxy that sits between the legit client and the server and forwards traffic live. The legit client does all the auth, all the heartbeats, all the token rotations, and the relay forwards the responses to N cracked clients.

To defeat relays, you need sequence counters (next section).


13. Step 7 — Fetching the gated data

This is where the actual protection lives. Every request here goes through multiple gates:

  1. Valid session exists
  2. Session is authenticated
  3. HWID matches
  4. Session hasn't aged out
  5. Last heartbeat is recent
  6. Heartbeat token matches the current one
  7. Sequence counter is strictly increasing
  8. User isn't banned, isn't expired, isn't blacklisted

Only then does the server construct the gated payload, encrypt it with a key derived from (session_key, hwid), and return it.

def handle_fetchdata(req, ip):
    app_data, app_id = resolve_app(req)
    if not app_data:
        return {"success": False, "message": "Invalid app"}

    # Find the encryption session (by either init session or user session)
    session = None
    for sid, s in _enc_sessions.items():
        if s.get("user_session") == req.sessionid or sid == req.sessionid:
            session = s
            break

    if not session:
        return {"success": False, "message": "No encryption session"}
    if not session["authenticated"]:
        return {"success": False, "message": "Not authenticated"}
    if session["hwid"] and not hwid_matches(session["hwid"], req.hwid or ""):
        return {"success": False, "message": "HWID mismatch"}
    if not session.get("last_hb"):
        return {"success": False, "message": "Heartbeat required first"}
    if time.time() - session["created"] > SESSION_MAX_AGE:
        return {"success": False, "message": "Session expired"}
    if time.time() - session["last_hb"] > HEARTBEAT_INTERVAL * 3:
        return {"success": False, "message": "Heartbeat stale"}

    # Heartbeat token check — constant time
    client_token = req.hb_token or ""
    server_token = session.get("hb_token") or ""
    if not client_token:
        return {"success": False, "message": "Missing heartbeat token"}
    if not hmac.compare_digest(client_token, server_token):
        return {"success": False, "message": "Stale heartbeat token"}

    # Sequence counter — must be strictly greater than the last seen value
    last_seq = session.get("client_seq", 0)
    if req.fetch_seq is not None:
        if req.fetch_seq <= last_seq:
            log_event(app_id, "security", "warning",
                      f"fetch_seq regression: {req.fetch_seq} <= {last_seq}")
            return {"success": False, "message": "Replay detected"}
        session["client_seq"] = req.fetch_seq
    elif app_data.get("require_fetch_seq"):
        return {"success": False, "message": "fetch_seq required"}

    # Database cross-check
    users = get_app_data(app_id, "users", [])
    user = next((u for u in users if u.get("session_id") == req.sessionid), None)
    if not user:
        return {"success": False, "message": "Invalid user session"}
    if user.get("banned"):
        return {"success": False, "message": "User banned"}

    # Blacklist checks
    blacklist = get_app_data(app_id, "blacklist", [])
    if is_blacklisted(blacklist, "hwid", req.hwid):
        return {"success": False, "message": "HWID blacklisted"}
    if is_blacklisted(blacklist, "ip", ip):
        return {"success": False, "message": "IP blacklisted"}

    # Construct the gated payload
    app_vars = get_app_data(app_id, "vars", {})
    gated = {
        "username": user["username"],
        "subscription": user["subscription"],
        "expiry": user["sub_expiry"],
        "level": user.get("user_vars", {}).get("level", "0"),
        "app_vars": app_vars,
        "app_name": app_data.get("name", ""),
        "version": app_data.get("version", "1.0.0"),
        "hwid": req.hwid,
        "nonce": secrets.token_hex(8),
        "ts": datetime.utcnow().isoformat() + "Z",
        "hb_token": session["hb_token"],
        "fetch_seq": session.get("fetch_count", 0) + 1,
    }
    session["fetch_count"] = session.get("fetch_count", 0) + 1

    # Encrypt with HWID-derived key
    encrypted = encrypt_for_hwid(session["enc_key"], req.hwid, gated)

    return {"success": True, "message": "Data retrieved", "encrypted_data": encrypted}
Enter fullscreen mode Exit fullscreen mode

Why per-HWID encryption

The critical line is this:

encrypted = encrypt_for_hwid(session["enc_key"], req.hwid, gated)
Enter fullscreen mode Exit fullscreen mode

The encryption key is derived from both the session key and the HWID. Two users with the same app have different session keys and different HWIDs, so their encrypted payloads are different and non-interchangeable.

Even if an attacker captures a full response and tries to use it on another machine:

# Client-side decryption, using the local machine's HWID
hwid_key = hmac.new(session_key, f"fetch-{local_hwid}", "sha256").digest()
plaintext = decrypt(captured_blob, hwid_key)  # → garbage
Enter fullscreen mode Exit fullscreen mode

The decryption produces garbage because the key derivation used a different HWID. The MAC check fails, the client refuses to use the data.

What the client does with the response

def load(self):
    if not self.decryption_key:
        return False

    next_seq = self._fetch_seq + 1
    r = self._req({
        "type": "fetchdata",
        "hwid": self._hwid(),
        "fetch_seq": next_seq,
    })
    if not r.get("success"):
        return False

    enc = r.get("encrypted_data")
    if not enc:
        return False

    # Derive the HWID-bound key
    session_key = base64.b64decode(self.decryption_key)
    hwid = self._hwid()
    ek = hmac.new(session_key[:32], f"fetchdata-enc-{hwid}".encode(), "sha256").digest()
    mk = hmac.new(session_key[:32], f"fetchdata-mac-{hwid}".encode(), "sha256").digest()

    # Verify MAC before decrypting (encrypt-then-MAC pattern)
    raw = base64.b64decode(enc)
    iv, ct, mac = raw[:16], raw[16:-32], raw[-32:]
    expected_mac = hmac.new(mk, iv + ct, "sha256").digest()
    if not hmac.compare_digest(mac, expected_mac):
        return False

    # Decrypt
    plaintext = hmac_stream(ek, iv, ct)
    plaintext = plaintext[:-plaintext[-1]]  # strip padding
    self._data = json.loads(plaintext)
    self._fetch_seq = next_seq
    return True
Enter fullscreen mode Exit fullscreen mode

Notice: MAC before decrypt. This is the encrypt-then-MAC pattern. If you decrypt first and check the MAC second, an attacker can submit malformed ciphertexts and learn things from how decryption fails (padding oracles, timing). Always verify the MAC first.


14. Step 8 — Anti-relay with sequence counters

A relay is the strongest attack against a session-based system. The attacker runs a proxy that connects to a legit licensed client on one side and to cracked copies on the other. The legit client does all the auth. The cracked copies just receive the decrypted data.

Standard session checks don't catch this because the relay is live. It's not replaying captured traffic, it's forwarding current traffic. The session is real, the heartbeat is real, the token is real.

The defense: give the client a strictly increasing counter that it must send with every fetchdata. The server tracks the highest value it's seen. Any request whose counter isn't strictly greater is rejected.

# Server side
last_seq = session.get("client_seq", 0)
if req.fetch_seq <= last_seq:
    log_event(app_id, "security", "warning", f"Replay: {req.fetch_seq} <= {last_seq}")
    return {"success": False, "message": "Stale fetch_seq"}
session["client_seq"] = req.fetch_seq
Enter fullscreen mode Exit fullscreen mode

Now consider a relay fronting N clones off one license. Each clone has its own local _fetch_seq counter. When clone #1 sends fetch_seq=5, the server records 5. Clone #2 sends fetch_seq=3 — rejected. Clone #2 retries with fetch_seq=6 — accepted, server now records 6. Clone #1's next call is fetch_seq=6 — rejected, it thinks the server is at 5.

Within a couple of requests, all but one clone is out of sync. The server logs a security warning every time.

Why not strictly require it?

The counter has to be added to the client SDK. If you turn on strict enforcement before all clients have updated, you break every old client. So make it opt-in per-app:

# App config
"require_fetch_seq": False  # or True once clients are updated

# Server check
if req.fetch_seq is not None:
    if req.fetch_seq <= last_seq:
        return reject()
    session["client_seq"] = req.fetch_seq
elif app_data.get("require_fetch_seq"):
    return {"success": False, "message": "fetch_seq required"}
Enter fullscreen mode Exit fullscreen mode

Start with soft mode (accept clients that don't send it, enforce for ones that do). Once you've rolled out the new SDK to your users, flip the app to strict mode. Old clients get rejected with a clear message.

The honest limitation

A patched client can just send fetch_seq incrementing integers regardless of what it's actually doing. The defense is cost — the attacker now has to reverse-engineer the counter semantics and get every clone to agree on a shared counter. It's not a wall, it's a speed bump. But it raises the cost of a relay from "run a proxy" to "write a coordinated multi-client sync layer," which is a much higher bar.


15. Step 9 — Binary integrity

The client's executable hashes itself and sends the hash with every request. The server compares against an allowlist the customer has curated. If it doesn't match, the server refuses.

# Client side
def binary_hash():
    """SHA-256 of the running executable."""
    import sys, os, hashlib

    if getattr(sys, "frozen", False):
        path = sys.executable  # PyInstaller / cx_Freeze
    elif __file__ and os.path.isfile(__file__):
        path = __file__  # Plain script
    else:
        return ""

    h = hashlib.sha256()
    with open(path, "rb") as f:
        for chunk in iter(lambda: f.read(65536), b""):
            h.update(chunk)
    return h.hexdigest()
Enter fullscreen mode Exit fullscreen mode
# Server side
def check_binary_hash(app_data, incoming_hash):
    approved = app_data.get("binary_hashes") or []
    if not incoming_hash:
        return False
    if not is_valid_sha256(incoming_hash):
        return False
    return incoming_hash.lower() in {h.lower() for h in approved}
Enter fullscreen mode Exit fullscreen mode

How the allowlist gets populated

Two modes:

Soft mode: log every observed hash but don't reject. The dashboard shows the customer "here are the hashes hitting your app — click the one that's your real build to approve it."

Strict mode: reject any hash that isn't approved. The customer flips this on once their allowlist is set.

# In the request handler
observed_ok = check_binary_hash(app_data, req.binary_hash)

if req.binary_hash:
    record_observed_hash(app_id, req.binary_hash, req.hwid, ip, observed_ok)

if app_data.get("binary_hash_required") and not observed_ok:
    log_event(app_id, "security", "warning", f"Binary hash rejected: {req.binary_hash[:16]}...")
    return {"success": False, "message": "Binary integrity check failed"}
Enter fullscreen mode Exit fullscreen mode

What this defeats

  • Modified executables. If an attacker patches the binary, the hash changes. Unless they also patch the hash function, they're rejected.
  • Repacked builds. Enigma, VMProtect, or custom packers change the binary. The resulting hash won't match.

What this doesn't defeat

  • Hash function patching. An attacker who replaces binary_hash() to always return the approved value gets around this. That's why you want to also do:
    • SDK version gating (next section)
    • Obfuscation so the hash function is hard to find
    • Multiple hash checks scattered through the binary so patching one doesn't help

The observed-hash buffer

A nice feature to include: keep a per-app circular buffer of recently-seen hashes with metadata (first seen, last seen, hit count, sample HWID). The dashboard shows this. When a customer ships a new build, they see the new hash appear in real time, click it, and it gets added to the allowlist. No more manual sha256sum workflows.

_observed_hashes = {}  # app_id -> {hash -> {first_seen, last_seen, count, hwid, approved}}

def record_observed_hash(app_id, h, hwid, ip, approved):
    if not is_valid_sha256(h):
        return
    h = h.lower()
    bucket = _observed_hashes.setdefault(app_id, {})
    entry = bucket.get(h)
    now = datetime.utcnow().isoformat()
    if entry is None:
        if len(bucket) >= 50:  # cap
            oldest = min(bucket, key=lambda k: bucket[k]["last_seen"])
            bucket.pop(oldest)
        bucket[h] = {
            "hash": h, "first_seen": now, "last_seen": now,
            "count": 1, "sample_hwid": hwid[:12], "approved": approved,
        }
    else:
        entry["last_seen"] = now
        entry["count"] += 1
        entry["approved"] = approved
Enter fullscreen mode Exit fullscreen mode

16. Step 10 — SDK version gating

Every time the customer ships an update to their app's SDK, any bypass written against the previous build stops working. SDK version gating is what forces clients to stay current.

The client sends its SDK version with every request:

SDK_VERSION = "1.3.0"

def _req(self, data):
    data["sdk_version"] = SDK_VERSION
    # ... rest
Enter fullscreen mode Exit fullscreen mode

The server compares against a per-app minimum:

def sdk_version_ok(client_ver, min_ver):
    if not min_ver:
        return True  # gate not enabled
    if not client_ver:
        return False

    def parse(v):
        return tuple(int("".join(c for c in p if c.isdigit()) or 0)
                     for p in str(v).split("."))

    try:
        return parse(client_ver) >= parse(min_ver)
    except Exception:
        return str(client_ver) >= str(min_ver)
Enter fullscreen mode Exit fullscreen mode

In the request handler:

_min_sdk = app_data.get("min_sdk_version")
if _min_sdk and not sdk_version_ok(req.sdk_version, _min_sdk):
    return {
        "success": False,
        "message": f"SDK update required. Minimum version: {_min_sdk}",
        "min_sdk_version": _min_sdk,
        "your_sdk_version": req.sdk_version or "unknown",
    }
Enter fullscreen mode Exit fullscreen mode

Why this works

Any patch written against SDK 1.2 has two options:

  1. Keep reporting sdk_version = "1.2.0" → server rejects (gate at 1.3)
  2. Report sdk_version = "1.3.0" → but the patch was written against 1.2's protocol, so any differences break

The attacker has to keep up with every SDK release. Each release is a cost they have to pay again.

Default off

Set min_sdk_version = None by default. Only flip it on after the customer has actually rolled out a new SDK to their users. Otherwise you lock out every existing client.


17. Step 11 — Rate limiting

Every public endpoint needs per-IP rate limiting, and login needs two — one per IP, one per account.

The implementation:

_rl_buckets = {}  # key -> [timestamps]

def rate_check(key, max_calls, window_seconds):
    now = time.time()
    cutoff = now - window_seconds
    bucket = _rl_buckets.setdefault(key, [])
    while bucket and bucket[0] < cutoff:
        bucket.pop(0)
    if len(bucket) >= max_calls:
        return False
    bucket.append(now)
    return True
Enter fullscreen mode Exit fullscreen mode

Usage as a FastAPI dependency:

def rate_limit(max_calls, window_seconds, scope=""):
    async def _dep(request: Request):
        ip = get_client_ip(request)
        key = f"{scope or request.url.path}:{ip}"
        if not rate_check(key, max_calls, window_seconds):
            raise HTTPException(429, "Too many requests")
    return _dep

@app.post("/api/login", dependencies=[Depends(rate_limit(10, 60, scope="login"))])
async def login_endpoint(req):
    ...
Enter fullscreen mode Exit fullscreen mode

The critical detail: default-reject unknown types

If your API has a type field that dispatches to different handlers, and you only apply rate limits for known types, an attacker can bypass rate limiting by sending an unknown type:

# BAD
limits = {"login": (20, 60), "register": (5, 60), "heartbeat": (240, 60)}
lim = limits.get(req.type)
if lim and not rate_check(...):
    return reject()
# Falls through to dispatch, then "unknown type"
Enter fullscreen mode Exit fullscreen mode

An attacker sends 10,000 requests with type = "xyz". None are rate-limited. All hit your dispatcher, all fail, but they've burned CPU.

Fix: apply a default limit for unknown types before dispatch:

lim = limits.get(req.type, (30, 60))  # default bucket
if not rate_check(f"sdk-{req.type}-ip:{ip}", *lim):
    return reject()
Enter fullscreen mode Exit fullscreen mode

Per-account throttling

IP-based limits don't help against distributed credential stuffing. An attacker with a botnet rotates through thousands of IPs, each making 1 request. Total: millions of attempts. Per-IP limit: never trips.

Per-account throttling fixes this:

if req.type == "login" and req.username:
    account_key = f"login:{app_id}:{req.username.lower()}"
    if not rate_check(account_key, 8, 300):
        return {"success": False, "message": "Too many attempts on this account"}
Enter fullscreen mode Exit fullscreen mode

Now the account can only be tried 8 times per 5 minutes, no matter how many IPs the attacker uses. Combined with per-IP limits, you've covered both attack modes.


18. Step 12 — Failed login tracking and auto-blacklist

Beyond rate limits, track failures and auto-blacklist after too many:

FAILED_LOGIN_LIMIT = 6
_failed_logins = {}  # (app_id, ip) -> {"count": int, "hwid": str, "last": float}

def record_failed_login(app_id, ip, hwid):
    key = (app_id, ip)
    rec = _failed_logins.get(key, {"count": 0, "hwid": "", "last": 0})
    rec["count"] += 1
    rec["last"] = time.time()
    if hwid:
        rec["hwid"] = hwid
    _failed_logins[key] = rec
    return rec["count"]

def maybe_auto_blacklist(app_id, ip):
    key = (app_id, ip)
    rec = _failed_logins.get(key)
    if not rec or rec["count"] < FAILED_LOGIN_LIMIT:
        return False

    blacklist = get_app_data(app_id, "blacklist", [])
    now = datetime.utcnow().isoformat() + "Z"

    if not is_blacklisted(blacklist, "ip", ip):
        blacklist.append({
            "id": secrets.token_hex(4),
            "type": "ip",
            "value": ip,
            "note": f"Auto-blacklisted after {rec['count']} failed logins",
            "added_at": now,
            "added_by": "system",
        })

    if rec.get("hwid") and not is_blacklisted(blacklist, "hwid", rec["hwid"]):
        blacklist.append({
            "id": secrets.token_hex(4),
            "type": "hwid",
            "value": rec["hwid"],
            "note": f"Auto-blacklisted after {rec['count']} failed logins",
            "added_at": now,
            "added_by": "system",
        })

    save_app_data(app_id, "blacklist", blacklist)
    del _failed_logins[key]
    return True
Enter fullscreen mode Exit fullscreen mode

Blacklisting both IP and HWID is important. If you only blacklist IP, an attacker on a dynamic IP gets a fresh start after a router restart. HWID blacklisting survives IP changes.

Notification

When auto-blacklist fires, notify the customer:

notify_customer(app_id,
    f"🚨 Auto-blacklisted\n\n"
    f"IP: {ip}\n"
    f"HWID: {rec['hwid'][:16]}...\n"
    f"After {rec['count']} failed logins\n"
    f"Last username tried: {req.username}")
Enter fullscreen mode Exit fullscreen mode

This is a critical part of the system. Without notifications, the customer has no idea they're being attacked. With them, they can manually escalate (ban the account, block a range) or let the auto-blacklist handle it.


19. Step 13 — Login anomaly detection

Beyond obvious attacks, there's the subtler stuff: legitimate credentials used from a suspicious location, an account that's suddenly being accessed from a new device, impossible travel between two logins.

None of these alone prove an account is compromised, but together they build a picture.

The simple scoring approach

Track each login and score it:

def score_login(user_id, ip, ua, geo):
    score = 0
    flags = []
    history = get_login_history(user_id)

    if not history:
        return {"score": 0, "flags": ["first_login"], "recommendation": "ok"}

    known_ips = {h["ip"] for h in history}
    known_countries = {h["country"] for h in history if h.get("country")}
    known_devices = {h["device"] for h in history if h.get("device")}

    # New IP
    if ip not in known_ips:
        score += 1
        flags.append("new_ip")

    # New country — higher weight
    country = geo.get("country", "Unknown")
    if country and country not in known_countries:
        score += 3
        flags.append("new_country")

    # New device
    device = parse_user_agent(ua)
    if device not in known_devices:
        score += 1
        flags.append("new_device")

    # Impossible travel
    if history and geo.get("lat") and geo.get("lon"):
        last = history[-1]
        if last.get("lat") and last.get("lon"):
            dist_km = haversine(
                last["lat"], last["lon"],
                geo["lat"], geo["lon"],
            )
            elapsed_hours = max((time.time() - last["ts"]) / 3600, 0.01)
            speed_kmh = dist_km / elapsed_hours
            # > 900 km/h is faster than a commercial jet
            if speed_kmh > 900 and dist_km > 100:
                score += 4
                flags.append(f"impossible_travel_{int(dist_km)}km")

    if score >= 6:
        return {"score": score, "flags": flags, "recommendation": "kill"}
    if score >= 3:
        return {"score": score, "flags": flags, "recommendation": "warn"}
    return {"score": score, "flags": flags, "recommendation": "ok"}
Enter fullscreen mode Exit fullscreen mode

Using the score

risk = score_login(user_id, ip, ua, geo)

if risk["recommendation"] == "kill":
    log_event(None, "security", "warning", f"Blocked login: {risk['flags']}")
    raise HTTPException(403, "Login blocked: suspicious activity detected")

if risk["recommendation"] == "warn":
    response["security_alert"] = {
        "message": "Unusual activity detected on your account",
        "flags": risk["flags"],
    }
Enter fullscreen mode Exit fullscreen mode

GeoIP lookup

For country/city/lat/lon, use a free service like ipwho.is:

def ip_geo(ip):
    if ip in ("127.0.0.1", "::1"):
        return {"country": "Local", "city": "localhost", "lat": 0, "lon": 0}
    try:
        with urllib.request.urlopen(f"https://ipwho.is/{ip}", timeout=2) as resp:
            data = json.loads(resp.read())
        if data.get("success"):
            return {
                "country": data.get("country", "Unknown"),
                "city": data.get("city", "Unknown"),
                "lat": float(data.get("latitude", 0)),
                "lon": float(data.get("longitude", 0)),
            }
    except Exception:
        pass
    return {}
Enter fullscreen mode Exit fullscreen mode

Cache the results for 24 hours per IP so you're not hitting the API on every request.

User agent parsing

Rough device fingerprint:

def parse_user_agent(ua):
    ua_lower = ua.lower()
    if "windows" in ua_lower: os_tag = "Windows"
    elif "mac" in ua_lower: os_tag = "macOS"
    elif "linux" in ua_lower: os_tag = "Linux"
    elif "android" in ua_lower: os_tag = "Android"
    elif "iphone" in ua_lower or "ipad" in ua_lower: os_tag = "iOS"
    else: os_tag = "Other"

    if "edg" in ua_lower: br = "Edge"
    elif "chrome" in ua_lower: br = "Chrome"
    elif "firefox" in ua_lower: br = "Firefox"
    elif "safari" in ua_lower: br = "Safari"
    else: br = "Other"

    return f"{br} on {os_tag}"
Enter fullscreen mode Exit fullscreen mode

It's crude but effective. "Firefox on Windows" vs "Chrome on Android" is enough to distinguish a normal login from a suspicious one.

Why this matters for licensing

For an app license system, anomaly detection catches:

  • Shared accounts (many logins from different locations)
  • Stolen sessions (sudden jump in location)
  • Account trading (multiple distinct device fingerprints on the same account)
  • Test-account abuse (a key marked for a demo suddenly being used by 100 different IPs)

None of these are security holes in the traditional sense, but they're the ways customers actually lose money.


20. Step 14 — The admin side

None of this matters if the customer can't see what's happening and take action. The admin dashboard needs:

Viewing

  • All apps owned by the customer
  • All license keys, sorted/filtered
  • All end-users, sorted/filtered
  • Active sessions, with kill buttons
  • The observed-hash buffer
  • Event log with severity filter

Actions

  • Generate new license keys (bulk)
  • Ban/unban end-users
  • Reset HWID for an end-user
  • Kill sessions (single or all)
  • Add to blacklist (IP, HWID, username)
  • Add/remove binary hashes from the allowlist
  • Toggle binary_hash_required, require_fetch_seq, etc.
  • Set min_sdk_version

The tier structure

If you're building this as a product, define plans:

PLANS = {
    "trial":      {"max_apps": 1,   "max_users_per_app": 100,   "price": 0},
    "solo":       {"max_apps": 3,   "max_users_per_app": 1000,  "price": 1900},
    "studio":     {"max_apps": 10,  "max_users_per_app": 10000, "price": 4900},
    "enterprise": {"max_apps": -1,  "max_users_per_app": -1,    "price": 14900},
}
Enter fullscreen mode Exit fullscreen mode

Enforce limits on app creation and key generation:

def assert_customer_can_create_app(principal):
    if is_admin(principal):
        return
    customer = principal["_customer"]
    limits = PLANS.get(customer["plan"], PLANS["trial"])
    max_apps = limits["max_apps"]
    if max_apps == -1:
        return
    current = len(customer_apps(customer["id"]))
    if current >= max_apps:
        raise HTTPException(402, f"App limit reached ({max_apps}). Upgrade to add more.")
Enter fullscreen mode Exit fullscreen mode

Notifications

The customer needs to know when something interesting happens:

  • New end-user registered
  • License key burned
  • End-user banned / auto-blacklisted
  • Concurrent sessions evicted (possible account sharing)
  • Binary hash rejected (tampered build)
  • Suspicious login on their own account

Route these through the customer's own notification channel (their Telegram bot, their Discord webhook, their email). Don't send them through the platform's channel — you want each customer's alerts to be private to them.


21. Step 15 — The client side

The client is the other half. Its responsibilities:

  1. Compute HWID
  2. Run init → auth → heartbeat → fetch in order
  3. Decrypt the fetched payload
  4. Periodically re-fetch, verify heartbeat
  5. Die cleanly-ish if any check fails

The skeleton

import requests, json, hashlib, hmac, base64, platform, socket, threading, time, sys, os

SDK_VERSION = "1.3.0"

class Auth:
    def __init__(self, api_url, app_name, owner_id):
        self.api = api_url
        self.name = app_name
        self.owner_id = owner_id
        self.session_id = None
        self.decryption_key = None
        self.hb_token = None
        self.fetch_seq = 0
        self.data = None
        self._alive = True
        self._hb_interval = 30

    def _hwid(self):
        raw = "|".join([
            platform.system(),
            platform.machine(),
            socket.gethostname(),
            platform.processor() or "",
        ]).encode()
        return hmac.new(b"hwid-salt", raw, hashlib.sha256).hexdigest()

    def _req(self, data):
        data["name"] = self.name
        data["ownerid"] = self.owner_id
        data["sdk_version"] = SDK_VERSION
        if self.session_id:
            data["sessionid"] = self.session_id
        if self.hb_token:
            data["hb_token"] = self.hb_token
        r = requests.post(self.api, json=data, timeout=10)
        return r.json()

    def init(self):
        r = self._req({"type": "init", "ver": "1.0.0"})
        if r.get("success"):
            self.session_id = r["sessionid"]
        return r

    def license(self, key):
        r = self._req({"type": "license", "key": key, "hwid": self._hwid()})
        if r.get("success"):
            self.session_id = r["sessionid"]
            self.decryption_key = r["decryption_key"]
            self._hb_interval = r.get("heartbeat_interval", 30)
        return r

    def load(self):
        """Fetch and decrypt gated data. Returns True only if everything works."""
        if not self.decryption_key:
            return False

        next_seq = self.fetch_seq + 1
        r = self._req({
            "type": "fetchdata",
            "hwid": self._hwid(),
            "fetch_seq": next_seq,
        })
        if not r.get("success"):
            return False

        enc = r.get("encrypted_data")
        if not enc:
            return False

        try:
            km = base64.b64decode(self.decryption_key)
            hwid = self._hwid()
            ek = hmac.new(km[:32], f"fetchdata-enc-{hwid}".encode(), hashlib.sha256).digest()
            mk = hmac.new(km[:32], f"fetchdata-mac-{hwid}".encode(), hashlib.sha256).digest()

            raw = base64.b64decode(enc)
            iv, ct, mac = raw[:16], raw[16:-32], raw[-32:]

            expected_mac = hmac.new(mk, iv + ct, hashlib.sha256).digest()
            if not hmac.compare_digest(mac, expected_mac):
                return False

            pt = hmac_stream(ek, iv, ct)
            pt = pt[:-pt[-1]]

            self.data = json.loads(pt.decode())
            self.fetch_seq = next_seq
            if self.data.get("hb_token"):
                self.hb_token = self.data["hb_token"]
            return True
        except Exception:
            self.data = None
            return False

    def _heartbeat_loop(self):
        fails = 0
        while self._alive:
            time.sleep(self._hb_interval)
            if not self._alive:
                break
            try:
                r = self._req({"type": "heartbeat", "hwid": self._hwid()})
                if r.get("success"):
                    fails = 0
                    if r.get("hb_token"):
                        self.hb_token = r["hb_token"]
                else:
                    fails += 1
            except Exception:
                fails += 1

            if fails >= 3:
                # Corrupt state and exit — NOT a clean exit that can be hooked
                self.data = None
                self.decryption_key = None
                self.session_id = None
                os._exit(0xFF)

    def start(self):
        threading.Thread(target=self._heartbeat_loop, daemon=True).start()

    def ready(self):
        return bool(self.data and self.data.get("username"))
Enter fullscreen mode Exit fullscreen mode

Usage

auth = Auth("https://api.example.com/api/1.2/", "my-app", "ownerid123")

if not auth.init().get("success"):
    sys.exit(1)

key = input("License key: ")
if not auth.license(key).get("success"):
    sys.exit(1)

auth.start()
time.sleep(auth._hb_interval + 1)  # wait for first heartbeat

if not auth.load():
    sys.exit(1)
if not auth.ready():
    sys.exit(1)

# App logic — all values come from the server
username = auth.data["username"]
max_level = auth.data["app_vars"].get("max_level", 10)
Enter fullscreen mode Exit fullscreen mode

The failure mode matters

Notice the os._exit(0xFF) on three heartbeat failures. Not sys.exit(). Not exit(0). os._exit() kills the process immediately without running cleanup handlers, without flushing buffers, without giving the app any chance to gracefully close.

Why? Because if the app exits cleanly, a patch can intercept the exit and prevent it. But if the app loses its state and then crashes, there's nothing to patch — the data is gone, the keys are gone, and the only path forward is a restart that re-runs the auth flow.

The more hostile the failure mode, the harder the app is to patch.


22. What's real defense and what's theater

After building this stuff for years, here's my honest assessment of each layer.

Real defenses

  • Server-gated data. The foundation. Everything else is built on it.
  • Proper password hashing (bcrypt/argon2). Non-negotiable. If your passwords aren't hashed with a slow algorithm and salt, nothing else matters.
  • Per-session random keys. Blast radius containment.
  • Constant-time comparison for tokens and HWIDs. Prevents timing oracles.
  • Heartbeat + rotating tokens. Kills replay. Real defense.
  • Per-account rate limiting. Defeats distributed credential stuffing. Real defense.
  • Session rotation on login. Kills stale sessions. Real defense.
  • Encrypt-then-MAC. Correct AEAD construction. Non-negotiable.
  • HWID binding (as a speed bump, not a wall). Slows attacks, doesn't stop them. But combined with everything else, useful.
  • Sequence counters. Real defense against relays, but defeatable by a patched client.

Theater (with a caveat)

  • Binary integrity checks. Defeated by patching the hash function. But: raises the bar, especially when combined with obfuscation and multi-point checks. Worth having.
  • SDK version gating. Defeated by a client that lies. But: forces the attacker to keep up with updates. Worth having.
  • Anti-debug checks. Trivially defeated by patching IsDebuggerPresent or using a debugger that hides itself. Pure theater on their own. Useful as speed bumps.
  • Packing / obfuscation (VMProtect, Enigma). Raises the cost of static analysis. Doesn't stop a determined attacker. Real value at scale.
  • Heartbeat required for fetchdata. Slows attacks that skip the heartbeat step. Theater if the client can just send a heartbeat first.

Things that feel secure but aren't

  • Encrypting passwords before hashing. Encryption is reversible — the key must live somewhere, so you've gained nothing. Hash, don't encrypt.
  • Adding "salt" by wrapping strings in symbols (##value@#). Not a cryptographic salt. Known format. Attacker can strip it.
  • Custom ciphers. Unless you have a formal security proof, you don't. Use AES-GCM.
  • Storing secrets in the client. Anything shipped to the client is compromised. Full stop.
  • Long expiry times. A 30-day session token is a 30-day window for a stolen token to be useful. Shorter is better.

The rule of thumb: if it makes the attacker's job harder, it's worth having. If it only makes you feel safer, it's theater. Most systems have both. The best systems know which is which.


23. Common mistakes

Every one of these has bitten someone.

1. Empty HWID skipping the check.

if stored_hwid and not hwid_matches(stored, incoming):
    reject()
Enter fullscreen mode Exit fullscreen mode

If incoming is empty, hwid_matches(stored, "") might return True (empty == empty or short-circuit). Free bypass. Always require non-empty.

2. Timing-attackable comparisons.

if client_token == server_token:
Enter fullscreen mode Exit fullscreen mode

Short-circuits on first differing byte. Use hmac.compare_digest.

3. The same message for "user not found" and "wrong password."

if not user:
    return "User not found"  # BAD — leaks usernames
if not verify_password(...):
    return "Wrong password"
Enter fullscreen mode Exit fullscreen mode

Always return "Invalid credentials."

4. Token/user info in URL parameters.

GET /fetchdata?token=abc123&user=alice
Enter fullscreen mode Exit fullscreen mode

Tokens end up in logs, browser history, referrer headers. Always POST, always in the body or a header.

5. Rolling your own crypto.
Yes, I'm repeating this. It's that important. Use AES-GCM, bcrypt, HMAC-SHA256. Standard primitives, standard constructions. If your code has a novel cipher, it's wrong.

6. Trusting the client to compute anything sensitive.
Anything the client computes can be recomputed wrongly. Anything the client stores can be extracted. The client's job is to talk to the server; the server's job is to decide.

7. Not rate-limiting by account, only by IP.
IP limits are trivially defeated by botnets. Account limits aren't.

8. Long-lived sessions.
A 30-day session means a stolen token is useful for 30 days. Refresh tokens should be short, access tokens shorter.

9. Storing secrets in plaintext on the server.
Telegram bot tokens, API keys, webhook secrets — all of them need at-rest encryption. A filesystem read shouldn't be enough to compromise everything.

10. No alerting on security events.
If you don't tell the customer when their app is being attacked, they can't respond. Notify on: auto-blacklist, concurrent session evictions, binary hash rejections, suspicious logins.

11. Assuming the attack isn't coming.
The day after you launch, someone will try. Assume it. Plan for it. Test it.


24. Testing your system

Before you ship it, test the attacks yourself. Every one of these should be blocked.

Session replay

  1. Log in as a user with a valid key
  2. Capture the exact fetchdata request and response
  3. Wait 60 seconds
  4. Replay the request
  5. Expected: rejected with "stale heartbeat token"

Session sharing

  1. Log in from machine A
  2. Copy the session ID and token to machine B
  3. Try to fetchdata from machine B
  4. Expected: rejected with "HWID mismatch"

Relayed clones

  1. Log in from one machine
  2. Run two processes that both call fetchdata with the same session
  3. Expected: one succeeds, one gets rejected with "stale fetch_seq"

Brute force from one IP

  1. Fire 100 login requests to the same account from one IP
  2. Expected: rate-limited after 8 requests; auto-blacklisted after 6 failures

Distributed brute force

  1. Fire 1 request per IP from many IPs (use a proxy service)
  2. Expected: account-level rate limit trips after 8 requests across all IPs

HWID bypass with empty field

  1. Send login with hwid: ""
  2. Expected: rejected with "HWID required"

Modified binary

  1. Patch the client executable
  2. Enable binary_hash_required on the app
  3. Log in with the patched binary
  4. Expected: rejected with "Binary integrity check failed"

Old SDK

  1. Set min_sdk_version = "1.3.0" on the app
  2. Send a request with sdk_version = "1.2.0"
  3. Expected: rejected with "SDK update required"

Token tampering

  1. Capture a valid hb_token
  2. Modify one character
  3. Send a request with the modified token
  4. Expected: rejected with "stale heartbeat token"

Password timing

  1. Time 10,000 requests with wrong passwords of varying prefix-lengths
  2. Expected: all requests take approximately the same time (no correlation with prefix match)

Run these before you have customers. Run them again after every significant change.


25. The honest ending

Everything is crackable.

Not "might be." Has been. Every commercial license system — Denuvo, BattlEye, KeyAuth, Auth.gg, iLok, Steam DRM — has been broken, repeatedly, by people who know what they're doing. There is no such thing as uncrackable software.

What you're building isn't a wall. It's a cost curve. The question isn't "can someone crack this?" — the answer to that is always yes. The question is:

  • How long does it take?
  • How much specialized skill does it require?
  • Does the crack survive your next update?

If cracking your app takes an afternoon, one person, and works forever — you've built nothing. If it takes weeks, requires deep reverse engineering skill, and breaks every time you ship a patch — you've built something real.

Not unbreakable. Just not worth it.

The layers in this post each raise the cost. Server-gated data makes patching impossible. HWID binding makes sharing detectable. Heartbeat tokens make replay useless. Sequence counters make relays desync. Binary integrity makes modification detectable. SDK versioning makes bypasses age out.

No single layer is a wall. Together they're a slope. The attacker has to keep climbing, and every step up costs them time and skill they'd rather spend elsewhere.

And that's the goal. Not perfection. Not certainty. Just a system expensive enough that the person attacking you decides their time is worth more somewhere else.

That's what real security looks like. Not paranoia, not theater — just an honest assessment of what each layer buys you, and the judgment to know when you've done enough.

Ship something simple first. Add layers only when you can justify them. And remember the foundation underneath everything:

You can't patch your way into data you don't have.

Everything else is commentary.

`

`

Top comments (1)

Collapse
 
dev_in_the_fog profile image
Jason Y. (dev_in_the_fog) •

Really thoughtful post! Documenting real-world engineering hurdles and actionable solutions like this brings immense value to the community.