DEV Community

jackymenCZ (jackymenCZ)
jackymenCZ (jackymenCZ)

Posted on

Agentix Lite v0.6: Building a Small Linux Security Sentinel That Knows When to Shut Up

There is a particular kind of software that sounds impressive until you ask a very simple question:

«“What happens when somebody tries to break the thing that is supposed to protect the thing?”»

That question is where Agentix Lite Sentinel started becoming interesting.

Agentix is a small, deterministic Linux host security agent designed around a rather unfashionable idea:

When the system is under pressure, the security agent should become less important than the application it is protecting.

No giant AI model.

No cloud analytics pipeline.

No 47 microservices arguing about who saw the packet first.

No infinite log queue.

And, importantly, no assumption that the machine has unlimited RAM because someone put “enterprise” in the README.

Agentix Lite Sentinel v0.6 is the result of several rounds of adversarial testing focused on one problem:

Can a security agent protect a Linux host without becoming another source of failure?

The answer is not “yes, absolutely, forever.”

The more honest answer is:

We built a bounded prototype, pushed its internal mechanisms hard, and reached the point where the next useful test is a real Linux VPS.

That distinction matters.


Part 1: What is Agentix?

Imagine a small security guard sitting next to a Linux server.

It watches a few things:

  • SSH activity
  • suspicious network activity
  • honeypot connections
  • requests hitting deliberately fake API endpoints
  • repeated probing patterns
  • system pressure
  • firewall actions

When something looks suspicious, it gives the source a reputation score.

For example:

Port scan +20
SSH brute force +40
Honeypot hit +100

These numbers are configurable. They are not magic.

Agentix then combines those signals with simple behavioral patterns.

For example:

scan
↓
pause
↓
SSH probing
↓
honeypot hit

That sequence is more interesting than a single request.

The important part is that Agentix does not need an AI model to understand it.

It is deterministic.

Same input.

Same state.

Same decision.

That makes it easier to test, reason about, and deliberately break.

And believe me, we tried to break it.


Why not just use logs?

Because logs are wonderful until there are too many of them.

A naive security script can easily turn into this:

internet
↓
events
↓
Python list
↓
more events
↓
larger Python list
↓
RAM
↓
OOM killer
↓
security agent has successfully defended the server
↓
by dying

That is not exactly the victory condition we wanted.

Agentix therefore treats memory as a budget.

Queues are bounded.

Firewall requests are bounded.

Ghost records are bounded.

Persistent actors are bounded.

Telemetry can be dropped.

That last part is especially important.

If the machine is overloaded, Agentix would rather lose some telemetry than block the application it is supposed to protect.

In other words:

«The security agent is allowed to forget. The web server is not allowed to wait for it.»


The Honey API

One of the more entertaining parts of Agentix is the Honey API.

It exposes fake endpoints that look interesting to automated scanners.

For example:

/api/v2/admin/config
/debug/env
/api/v2/user/update/{user_id}

These endpoints are deliberately attractive to scanners.

But the Honey API does not need to store everything an attacker sends.

That would be a fantastic way to build a very expensive scrapbook of garbage.

Instead, Agentix extracts the small amount of information it needs for classification and keeps telemetry bounded.

The actual request body and raw headers are not treated as permanent security data.

The idea is simple:

Give scanners something interesting enough to touch, then record the fact that they touched it.

Not their entire autobiography.


“What if the attacker floods the security agent?”

This became one of the main design questions.

Suppose 50,000 firewall actions are generated.

A bad implementation might create:

50,000 objects
50,000 futures
50,000 callbacks
50,000 subprocesses

and then politely wait for Linux to kill it.

Agentix instead has a bounded firewall queue.

The V6 stress test submitted:

50,000 firewall requests

The queue stayed at:

512

The remaining requests were shed.

That sounds strange at first.

Shouldn't a security system process everything?

No.

Not if processing everything means killing itself.

A bounded security system has to be willing to say:

«“I have 512 seats. The other 49,488 requests are not getting in.”»

That is not a detection failure.

It is controlled degradation.


What happens to the dropped firewall actions?

This exposed another interesting problem.

Imagine the firewall queue is full and thousands of additional actions are rejected.

If Agentix writes one log entry for every rejected action, we have just created another denial-of-service mechanism.

So V6 aggregates the shedding information.

Instead of:

DROP 1.2.3.4
DROP 1.2.3.5
DROP 1.2.3.6
...

the system can report the important fact:

firewall_shed = true

with counters and aggregated state.

The administrator learns:

“The firewall actuator is under pressure.”

They do not need to read 49,488 nearly identical log messages while drinking their morning coffee.

Coffee is already complicated enough.


IPv6 changed the ghost problem

One of the more interesting problems appeared when thinking about IPv6.

With IPv4, an actor can be represented by:

203.0.113.42

Simple.

But IPv6 gives attackers enormous address space.

A scanner can rotate addresses inside the same "/64".

If the security system treats every address as a completely independent actor, it can end up doing this:

IPv6 A → actor
IPv6 B → actor
IPv6 C → actor
IPv6 D → actor
...

Eventually the machine spends more effort remembering the attacker than detecting the attacker.

V6 therefore introduces IPv6 ghost aggregation.

For the relevant ghost identity, addresses inside the same "/64" can share a compact identity.

The system does not store the original IP inside the ghost record.

The result is bounded memory behavior rather than endless IPv6 identity churn.

A stress test with:

10,000 IPv6 churn events

ended with:

512 ghosts retained

And a separate test with:

500 IPv6 addresses
inside one /64

produced:

1 ghost identity

This is one of those cases where IPv6 politely reminds you that “unique IP address” and “unique human attacker” are not the same concept.


SQLite is part of the security model

Agentix uses SQLite because the project is deliberately small.

No external database cluster.

No Redis dependency.

No Kafka.

No Kubernetes operator whose sole purpose is to restart another Kubernetes operator.

SQLite is enough for the compact persistent state we need.

But SQLite has a property that matters:

WAL mode introduces a checkpointing problem.

Writes go into the WAL.

Eventually the WAL needs to be checkpointed.

And under sustained activity, that work itself can become part of the performance problem.

So V6 moved routine checkpoint work into a separate maintenance worker.

The main event-processing path does not deliberately perform the expensive checkpoint operation.

The database also has explicit storage budgets.

When storage pressure becomes dangerous, Agentix can shed telemetry.

Again:

«Lose data before losing the machine.»


The interesting bug we found

During V6 testing, we discovered something important.

A WAL checkpoint could successfully recycle pages without necessarily making the physical WAL file look small.

So a naive measurement such as:

WAL exists

could give a misleading impression about actual storage pressure.

The V6 maintenance path therefore pays attention to actual checkpoint progress and can perform a "TRUNCATE" after appropriate successful checkpoint conditions.

The resulting stress test:

5,000 SQLite writes

finished with:

WAL = 0 bytes

under the synthetic hard-guard scenario.

This is exactly why I prefer testing over declarations.

A README can say “WAL is handled.”

A stress test can say:

«“No, this particular assumption was incomplete.”»

The stress test wins.


What Agentix does NOT claim

This is probably the most important section.

Agentix Lite Sentinel v0.6 is not being presented as:

  • a DDoS mitigation service
  • a commercial WAF
  • a carrier-grade firewall
  • an AI SOC
  • an intrusion-prevention system proven against real-world attacks
  • a replacement for professional infrastructure security
  • a system proven to survive arbitrary hostile traffic

The tests were performed in a controlled environment.

They demonstrate bounded behavior under specific synthetic workloads.

They do not prove what happens after seven days on a public VPS.

That is the next experiment.

And honestly, that's the fun part.


Part 2: For the programmers

Now let's take the hood off.

The core Agentix architecture is intentionally small.

Conceptually:

                Internet
                   |
          +--------+--------+
          |                 |
      Honey API         TCP honeypot
          |                 |
          +--------+--------+
                   |
          Unix datagram lanes
                   |
          bounded admission
                   |
          bounded processing
                   |
        +----------+----------+
        |                     |
    patterns              reputation
        |                     |
        +----------+----------+
                   |
                SQLite
                WAL
                   |
         bounded firewall queue
                   |
              batched nft
Enter fullscreen mode Exit fullscreen mode

The implementation is primarily Python.

The Honey API uses FastAPI.

The deployment environment uses:

  • systemd
  • Docker
  • nftables
  • SQLite
  • Unix datagram sockets

There is deliberately no external LLM dependency in the detection path.


Three ingress lanes

Agentix uses separate Unix datagram paths for different classes of telemetry.

Conceptually:

normal
honey-critical
host-critical

They have independent bounded queues and admission state.

The receiver does not blindly trust a client-supplied priority field.

The producer determines the lane.

The receiver validates the source.

This matters because otherwise an attacker might simply say:

{
"priority": "critical"
}

and congratulations, everyone is important now.


Non-blocking telemetry

The telemetry client is intentionally boring.

Very boring.

It performs one non-blocking "sendto()" attempt.

If the socket cannot accept the datagram:

EAGAIN
EWOULDBLOCK
ENOBUFS
ENOENT
ECONNREFUSED
EPIPE

the event is dropped.

There is:

no retry loop
no sleep
no disk spool
no hidden queue
no asyncio task

This is intentional.

The protected application should not wait for Agentix.

A security agent should not turn this:

HTTP request → 20 ms

into this:

HTTP request
↓
security telemetry
↓
security telemetry retry
↓
security telemetry retry
↓
security telemetry queue
↓
HTTP request → timeout

The request wins.


Firewall actuator

The firewall layer is bounded at multiple points.

The request queue has a maximum size.

Requests are deduplicated.

Low-priority requests can be shed.

The actuator batches addresses and uses a single nftables transaction instead of spawning one subprocess per IP.

Conceptually:

request
↓
deduplicate
↓
priority queue
↓
bounded mailbox
↓
batch
↓
nft -f

The completion mailbox is also bounded.

This matters because bounding only the input queue is not enough.

A system can have:

bounded input
+

unbounded results

still eventually dead

Every queue is a potential memory leak wearing a data structure costume.


SQLite storage model

The persistent state is intentionally compact.

Actor records contain information such as:

IP
first_seen
last_seen
score
attempt_count
pattern
action

Evicted actors can leave behind a compact HMAC-based ghost representation.

The original IP does not have to remain in the ghost record.

The V6 IPv6 logic adds aggregation at the identity layer so that a rotating "/64" does not necessarily create an unlimited number of persistent identities.

The important design constraint is:

attacker-controlled cardinality
↓
bounded representation

That pattern appears throughout the project.


WAL maintenance

The main SQLite connection has automatic checkpoint behavior disabled.

A separate maintenance connection performs checkpoint work.

The maintenance worker watches:

.db
-wal
-shm

as a combined footprint.

V6 additionally tracks checkpoint behavior rather than assuming:

checkpoint_called == checkpoint_completed

That distinction is important under load.

A "PASSIVE" checkpoint may make progress without being able to finish everything immediately.

Therefore V6 uses bounded maintenance behavior rather than blocking the event path waiting for SQLite to become perfectly calm.

If storage becomes unhealthy, telemetry can be shed.


Ghost admission

Ghost storage is also bounded.

V6 adds explicit admission control so that an attacker cannot turn ghost creation into an unlimited side channel.

This is particularly relevant for IPv6.

The basic philosophy is:

high-confidence actor
↓
persistent state

low-confidence / short-lived actor
↓
maybe nothing

evicted actor
↓
compact ghost

The system should not remember everything.

It should remember what is useful.

That sounds obvious until you build a system that receives millions of unique inputs.


Resource limits

The deployment includes system-level limits as a second line of defense.

The V6.0.1 profile uses approximately:

MemoryHigh 160 MiB
MemoryMax 180 MiB
CPUQuota 50%
TasksMax 32
LimitNOFILE 4096

The earlier benchmark produced approximately:

RSS ~135 MiB
Python heap ~10–13 MiB

depending on workload and environment.

This distinction matters.

Python heap is not process RSS.

The interpreter, SQLite, native libraries, allocator behavior and other runtime components all contribute to RSS.

So saying:

«“Agentix uses 10 MB”»

would be technically misleading.

The benchmark actually showed roughly 135 MiB process RSS.

That number is environment-dependent.


The V6 stress results

Here is the more useful table.

Test| Result
Python regression suite| 52/52 PASS
Pattern workload| ~4,284 events/sec
Health workload| ~9,622 events/sec
100k persistent actors| completed
IPv6 churn| 512 ghosts retained
500 IPv6 addresses /64| 1 ghost
Firewall submissions| 50,000
Firewall queue maximum| 512
Transport flood| 10,000 datagrams
Transport queue maximum| 64
SQLite stress| 5,000 writes
Final WAL in stress| 0 bytes

These are test observations, not capacity guarantees.

Especially the events/sec numbers.

They were measured in one environment.

A VPS with different CPU, storage and kernel behavior will produce different numbers.


The 0.6.1 deployment hardening

After V6, I made a small deployment-only revision.

Agentix Lite Sentinel v0.6.1 does not change the detection architecture.

It tightens deployment behavior.

Among other things:

  • Python 3.12+ is required by the installer
  • systemd resource limits are explicit
  • filesystem protection is enabled
  • "CAP_NET_ADMIN" is isolated to the service
  • "NoNewPrivileges" is enabled
  • write access is restricted to Agentix runtime/data paths
  • reverse-proxy IP handling is more careful about trusted proxy hops
  • deployment documentation includes the Shadow Mode procedure

The important distinction is:

v0.6 = architecture
v0.6.1 = deployment hardening


What happens next?

This is where the project gets much more interesting.

Not V7.

Not another 40-page architecture diagram.

A VPS.

The first deployment should run in:

Shadow Mode

with enforcement disabled.

The goal is not to block attackers yet.

The goal is to observe reality.

For approximately seven days, I would record:

actors
ghosts
SQLite size
WAL size
checkpoint progress
storage pressure
firewall shedding
firewall actions
transport drops
RSS
CPU
service restarts

And then compare Agentix's observations against the actual Nginx/Caddy/application logs.


The questions the internet gets to answer

There are several questions we simply cannot answer from a laptop benchmark.

  1. How fast does actor churn happen?

Synthetic tests can generate 100,000 actors.

Real internet traffic has different distributions.

Bots repeat.

Bots rotate.

Some scan slowly.

Some hit everything at once.

Some behave strangely enough to make your pattern detector question your life choices.


  1. Does IPv6 aggregation behave well in reality?

The "/64" assumption is useful, but the real question is:

What does actual hostile IPv6 traffic look like on this particular VPS?

That is an empirical question.


  1. How does SQLite behave on cheap VPS storage?

A local development machine and a budget VPS can have very different I/O behavior.

The seven-day run should tell us whether checkpoint scheduling remains comfortably inside the resource budget.


  1. How often does firewall shedding actually happen?

This is probably one of the most valuable real-world measurements.

If:

firewall_shed = 0

for seven days, excellent.

If it happens occasionally during bursts, that's useful.

If it happens constantly, the budget or actuator design needs attention.

We shouldn't guess which one will happen.


And then there is the most important metric

Not:

AI accuracy

Not:

tokens saved

Not:

number of lines of code

Not:

how futuristic the architecture sounds

The useful question is:

«Does Agentix observe hostile activity without becoming a problem itself?»

That is the experiment.


Final thoughts

Agentix Lite Sentinel started as a relatively simple idea:

watch Linux
↓
recognize suspicious behavior
↓
remember enough
↓
react carefully

The interesting engineering turned out not to be the detection rules.

It was the boundaries.

How much memory can an attacker indirectly make us consume?

How many events can enter?

How many firewall operations can wait?

What happens when SQLite is busy?

What happens when IPv6 gives us an absurd number of addresses?

What happens when the firewall cannot keep up?

What happens when telemetry disappears?

And perhaps the most important one:

«What happens when Agentix itself is the thing under pressure?»

The answer we built around is deliberately conservative:

bounded memory
bounded queues
bounded storage
bounded firewall work
non-blocking telemetry
deterministic decisions
controlled shedding

If something has to be sacrificed, Agentix sacrifices telemetry before it sacrifices the application.

That's not glamorous.

It's engineering.

And now the interesting part begins.

The next version of Agentix should not be invented in an IDE.

It should be written by the internet.

Preferably without setting the VPS on fire.

If you're a Linux, networking, SQLite, nftables, Python, or infrastructure engineer and see a flaw in the design, I'd genuinely like to hear about it.

Especially if you can reproduce it.

Because at this point, another theoretical attack diagram is worth less than one angry little bug report from a real Debian machine.

Honey endpoint is waiting. 🍯

Bring your scanner.

Top comments (0)