DEV Community

Cover image for SwarmReconGuard: Detecting Distributed Agent Reconnaissance
mech.app
mech.app

Posted on Originally published at mech.app

SwarmReconGuard: Detecting Distributed Agent Reconnaissance

Autonomous agents create a new reconnaissance threat model. A single adversary can distribute system mapping across hundreds or thousands of identities so each request remains valid, authenticated, and low-rate while the population collectively builds a complete API map. Traditional defenses fail because no individual agent violates rate limits or exhibits suspicious behavior.

SwarmReconGuard formalizes this as Distributed Collective Reconnaissance (DCR) and presents a black-box detection benchmark. The defender sees only service-boundary telemetry: request timestamps, endpoints, response codes, and latencies. No access to agent coordination channels, internal state, or cross-identity correlation at the application layer.

The Threat Model

DCR works because modern API infrastructure is designed to handle multi-tenant traffic. Each agent:

  • Uses a unique, valid credential
  • Stays well below per-identity rate limits
  • Issues semantically reasonable requests
  • Follows normal retry and backoff patterns
  • Exhibits human-like timing jitter

The swarm collectively:

  • Enumerates all endpoints and parameters
  • Probes authorization boundaries
  • Maps data relationships
  • Identifies rate limit thresholds
  • Discovers error-handling behavior

Traditional defenses (per-IP rate limiting, WAF rules, anomaly detection on single identities) see nothing unusual. Each agent looks like a legitimate low-volume user.

Detection Architecture

SwarmReconGuard operates at the service boundary with only observable telemetry:

Input signals:

  • Request timestamp
  • Endpoint path
  • HTTP method
  • Response status code
  • Response latency
  • Identity token (opaque to detector)

Detection pipeline:

  1. Temporal aggregation: Sliding windows (1min, 5min, 15min) compute population-level statistics
  2. Feature extraction: Endpoint diversity, request rate variance, coverage velocity, error distribution
  3. Statistical testing: Compare observed population behavior against benign baseline distributions
  4. Threshold logic: Flag when collective patterns exceed learned bounds

The detector never sees what individual agents are doing. It only observes that the population as a whole is covering the API surface faster or more systematically than legitimate usage patterns.

Benchmark Design

The reproducible study uses Docker isolation to ensure telemetry integrity:

Test matrix:

  • 10 to 10,000 virtual identities
  • 11 distinct behaviors (benign browsing, systematic enumeration, adaptive probing)
  • 440 test runs
  • 3,666,300 total requests

Benign baselines:

  • Random endpoint selection
  • Zipf-distributed access (mimicking real user patterns)
  • Session-based browsing with temporal locality

Attack policies:

  • Breadth-first enumeration
  • Depth-first parameter exploration
  • Adaptive probing (adjust based on responses)
  • Hybrid strategies

Each run produces complete telemetry logs with ground-truth labels for evaluation.

Detection Methods Compared

Detector Detection Rate (Known) Detection Rate (Unseen) False Positive Rate Scale Sensitivity
Gaussian likelihood-ratio 100% 3% 0% Low
CUSUM (basic) 36.1% 36.1% 1.25% Medium
Hybrid CUSUM 85.7% Not reported 0% High
Semantic embedding 67% 12% 2.1% Low
Graph-based 54% 8% 0.5% Medium
Kernel density 48% 15% 3.2% Low

The Gaussian detector achieves perfect accuracy on known attack policies but fails catastrophically on unseen strategies (3% detection). This exposes the policy-generalization gap: detectors trained on specific reconnaissance patterns miss novel approaches.

Hybrid CUSUM performs best at scale (10,000 identities) by combining change-point detection with population-level feature tracking. It flags when the collective endpoint coverage rate or error distribution shifts beyond baseline variance.

Implementation Considerations

Real-time constraints:

Sliding-window aggregation must complete within the window duration. For 1-minute windows at 10,000 requests/second, you need sub-100ms feature extraction and scoring.

State management:

The detector maintains:

  • Per-identity request counters (sparse, most identities are inactive)
  • Population-level histograms (endpoint coverage, status codes)
  • Baseline distributions (learned from benign traffic)
  • CUSUM accumulators (one per monitored feature)

At 10,000 identities, this requires roughly 50MB of hot state if you track the last 15 minutes of activity.

False positive handling:

Legitimate traffic patterns shift during:

  • Product launches (new endpoints see coordinated exploration)
  • Documentation updates (users try new API features)
  • Integration testing (QA teams systematically probe endpoints)

You need exposure-aware thresholds that account for expected coverage velocity during these events.

Observability:

Log every detection event with:

  • Triggering feature (endpoint diversity, coverage rate, error spike)
  • Population size at trigger time
  • Top contributing identities (for incident response)
  • Baseline distribution snapshot

This lets you tune thresholds and investigate false positives without re-running the entire detection pipeline.

Deployment Shape

class SwarmDetector:
    def __init__(self, window_minutes=5, threshold_sigma=3.5):
        self.window = window_minutes * 60
        self.threshold = threshold_sigma
        self.baseline = BaselineDistribution()
        self.cusum = CUSUMAccumulator()

    def process_request(self, timestamp, identity, endpoint, status):
        # Update population statistics
        self.update_coverage(timestamp, endpoint)
        self.update_error_rate(timestamp, status)

        # Compute features every N requests
        if self.should_evaluate(timestamp):
            features = self.extract_features()
            score = self.cusum.update(features, self.baseline)

            if score > self.threshold:
                return DetectionEvent(
                    timestamp=timestamp,
                    score=score,
                    features=features,
                    population_size=self.active_identities()
                )
        return None

    def extract_features(self):
        return {
            'endpoint_diversity': self.unique_endpoints() / self.total_requests(),
            'coverage_velocity': self.new_endpoints_per_minute(),
            'error_concentration': self.error_gini_coefficient(),
            'request_rate_variance': self.population_rate_stddev()
        }
Enter fullscreen mode Exit fullscreen mode

The detector runs inline in your API gateway or as a sidecar that consumes request logs from a stream (Kafka, Kinesis). It does not block requests. Detection events trigger alerts or automated responses (rate limit the population, require additional auth challenges, flag for manual review).

Failure Modes

Slow reconnaissance:

If the swarm spreads activity over days or weeks, population-level features stay within baseline variance. The detector sees gradual coverage growth that looks like organic user adoption.

Mitigation: Track long-term coverage trends and flag when total API surface knowledge exceeds expected user behavior (most users only touch 5-10% of endpoints).

Coordinated benign activity:

A new mobile app release causes thousands of users to explore the same new features simultaneously. The detector sees rapid endpoint coverage and flags a false positive.

Mitigation: Integrate with deployment events and suppress alerts during known rollout windows.

Feature drift:

Baseline distributions shift as your API evolves (new endpoints, deprecated routes, changing usage patterns). The detector's learned baseline becomes stale.

Mitigation: Continuous baseline retraining with exponential decay weighting (recent traffic matters more than old patterns).

Adaptive adversaries:

Once attackers know you detect coverage velocity, they slow down or mimic Zipf distributions. The cat-and-mouse game continues.

Mitigation: Combine multiple detection features (coverage, error patterns, timing correlations) so evading one signal doesn't defeat the entire system.

Technical Verdict

Use SwarmReconGuard-style detection when:

  • You operate a large API surface (100+ endpoints) with many identities
  • Individual rate limiting and WAF rules are insufficient
  • You need black-box detection without application-layer correlation
  • You can tolerate 1-3% false positives during normal operation
  • You have real-time telemetry infrastructure (log streaming, aggregation)

Avoid or defer when:

  • Your API has fewer than 20 endpoints (easier to monitor manually)
  • You have strong application-layer identity correlation (can track cross-identity patterns directly)
  • Your traffic volume is too low for statistical detection (fewer than 1,000 requests/hour)
  • You lack baseline data (need weeks of benign traffic to train distributions)
  • Your API usage patterns shift daily (makes baseline learning impractical)

The policy-generalization gap is the key limitation. Detectors trained on known reconnaissance strategies miss novel approaches. You need continuous evaluation against new attack policies and hybrid detection that combines multiple statistical signals. This is not a deploy-and-forget solution. It requires ongoing tuning, baseline updates, and integration with incident response workflows.

Source Links

Top comments (0)