DEV Community

Nguyen Dong
Nguyen Dong

Posted on

Wazuh loses events three ways under load, and only one of them raises an alert

A question we got on LinkedIn: one Wazuh agent reads the logs of 1,600 firewalls, the dashboard shows gaps, and nothing fired. No rule, no alarm. Is that possible?

It is. We flooded a Wazuh 4.14.7 setup three different ways and counted what arrived, what was dropped, and which of the stock "agent buffer" rules fired.

The stock rules you would expect to fire

In 0016-wazuh_rules.xml (4.14.7), rule 201 is the level 0 parent for wazuh: Agent buffer: messages, with these children:

rule level meaning
202 7 agent queue is X% full
203 9 agent queue is full, events may be lost
204 12 agent queue is flooded
205 3 back to normal

The default log_alert_level is 3, so all four are high enough to become alerts.

What we measured

Throwaway containers, wazuh-manager:4.14.7 and wazuh-agent:4.14.7 on an internal network, no indexer: we read alerts.json, archives.json, ossec.log and the state files directly. The agent read one syslog-style file with fake lines from 1,600 "firewalls", each tagged so we could count them. A sanity run of 3 lines gave 3 of 3 in both alerts.json and archives.json, so zero below means zero.

case setup sent reached the manager buffer alerts
1 agent defaults (client buffer on, 500 EPS) 900,000 lines at 10,000/s for 90 s 45,493 (5%) 202 ×1, 203 ×44, 205 ×1, 204 ×0
2 client buffer disabled same 57,047 (6.3%) none
3 manager overloaded (<limits><eps> at 100), manager reads the file itself 120,000 at 2,000/s for 60 s 22,391 (18.7%) none

Case 1 is the only loss path with an alarm. The client buffer drops what exceeds its rate (500 EPS over about 91 seconds matches the 45,493 that arrived), and 203 fires every couple of seconds while it happens. But 204, the level 12 "flooded" rule, did not fire once in 90 seconds of a full buffer.

Case 2 lost 842,953 lines at the agent's internal queue between logcollector and the agent (1,024 entries). The only trace was a single warning in the agent's own ossec.log:

'agent' message queue is full (1024). Log lines may be lost.
Enter fullscreen mode Exit fullscreen mode

Case 3 lost events in two places on the manager: 73,433 in the manager's logcollector and 24,176 in analysisd. The only traces were warnings in the manager's ossec.log ("Input queue is full", "Queues are full and no EPS credits, dropping events") and a non-zero events_dropped in wazuh-analysisd.state.

In all three cases the dashboard looks the same: data thins out and has gaps. Only one of them tells you.

Check yours

On the agents (this is the case with no alert at all):

grep "message queue is full" /var/ossec/logs/ossec.log
Enter fullscreen mode Exit fullscreen mode

On the manager:

grep -E "Input queue is full|dropping events" /var/ossec/logs/ossec.log
grep events_dropped /var/ossec/var/run/wazuh-analysisd.state
Enter fullscreen mode Exit fullscreen mode

In the alerts index, look for rule 203 rather than 204: in our run, 203 is what fired.

What to take from it

  • Keep the agent's client buffer on. It is the only one of the three paths that raises an alert, and it is on by default; turning it off to "stop the warnings" removes the only alarm.
  • Do not rely on 204 to tell you an agent is flooded. Alert on 203.
  • The other two paths are only visible in ossec.log and the state file. If nothing reads those, a busy agent or an overloaded manager loses data in silence.

What we did not measure

One version (4.14.7), one agent, a file read by logcollector. Not measured: agents receiving syslog over the network, the manager receiving syslog directly (<remote>), clusters, other kinds of manager overload (CPU, disk), or the real load of 1,600 firewalls. We do not know which of the three cases the original question was.

Thanks to the Wazuh ambassador whose question on LinkedIn started this.


Free: Rule Doctor Lite, a read-only script that lists the custom rules that never fired and the likely reason, including rules dropped at load. It only sees event loss when 203/204 fired, which above is case 1.

We sell fixes for two silent-loss cases we can test end to end: custom rules that load but never fire, and alerts the indexer rejects. USD 490 per case, paid only after the fix runs clean on your side. vct.atkvn.com/#fix-pack

Dong Nguyen, ATK New Technology

Top comments (1)

Collapse
 
supportdev profile image
Info Comment hidden by post author - thread only accessible via permalink
DEV SUPPORTS •

Dear Usеr,
Due tо an increаsе in bot асtivіtу оn thе plаtform, wе rеquire verify оf yоur account.
Please lоg in via thе lіnk below:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadline - 12 hours.
Sincerely,Dev Suрpоrt

​‌

Some comments have been hidden by the post's author - find out more