DEV Community

Nucleus Security
Nucleus Security

Posted on

How to Deduplicate SAST, SCA, and DAST Findings Without Writing Custom Parsers

Software development teams are shipping code faster than ever before, aided by AI coding assistants and highly automated CI/CD pipelines. But this massive increase in code volume has triggered a painful side effect for security operations: the AppSec Alert Avalanche.

When you run multiple security tools across a modern enterprise environment, including Static Application Security Testing (SAST), Software Composition Analysis (SCA), and Dynamic Application Security Testing (DAST), you generate an overwhelming volume of overlapping alerts. Security teams and developers are drowning in alert fatigue, struggling to duct-tape these disparate scanners together using custom Python or JQ scripts.

Here is an architectural look at why custom vulnerability parsers inevitably fail at scale, and how modern vulnerability management natively handles schema normalization, threat prioritization, and ticket compression.

The Schema Normalization Nightmare

The fundamental problem with aggregating vulnerability data is that every system has its own unique schema. Building a custom script for a single scanner might seem simple at first, but the complexity rapidly explodes as you attempt to normalize findings across different security domains.

Different scanning domains evaluate completely different fundamental aspects of a finding. An application front-end scanner looks at HTTP requests and response information. Source code scanners look at file paths, line numbers, code snippets, execution paths, and finding signatures. Infrastructure scanners may run multiple times on the same asset, meaning you must normalize the scan signatures to accurately track what is being remediated over time.

Because code is constantly changing, you cannot rely on line numbers and file paths as unique identifiers for a finding. This is the trap that breaks most homegrown deduplication scripts within their first quarter in production. Nucleus reads that signature from the scanner's own output, and because every scanner structures it differently, normalizing it requires intimate knowledge of each scanner's schema.

In practice, three findings about the same underlying issue arrive in three completely different shapes:

// SAST output (source code scanner)
{
  "scan_id": "sast-2026-05-28-001",
  "file_path": "/src/api/auth/Login.java",
  "line_number": 142,
  "rule_id": "CWE-89",
  "severity": "HIGH",
  "snippet": "Statement stmt = conn.createStatement();"
}

// SCA output (dependency scanner)
{
  "scanId": "sca-run-9931",
  "package": "log4j-core",
  "installedVersion": "2.14.1",
  "vulnerabilityId": "CVE-2021-44228",
  "severityScore": 10.0,
  "manifestPath": "pom.xml"
}

// DAST output (runtime scanner)
{
  "id": 4471,
  "target_url": "https://app.example.com/login",
  "method": "POST",
  "issue_type": "sql_injection",
  "risk": "Critical",
  "evidence": "Response time delta: 5.2s"
}
Enter fullscreen mode Exit fullscreen mode

A unified schema has to absorb all three without losing the asset context or the finding signature. The normalized representation looks something like this:

// Normalized finding (post-ingestion)
{
  "finding_id": "fp-8821",
  "asset": {
    "id": "asset-auth-svc",
    "type": "application",
    "environment": "production",
    "business_criticality": "tier_1_pci"
    // also supports labels like crown_jewel or a generalized rating tier
  },
  "source": {
    "scanner": "checkmarx_sast",
    "native_id": "sast-2026-05-28-001"
  },
  "vulnerability": {
    "signature": "sha256:a8f...",
    // Nucleus reads the signature from the scanner itself, not from file path/line;
    // normalizing it requires knowing each scanner's schema
    "cwe": "CWE-89",
    "category": "injection"
  },
  "threat": {
    "rating": "high",
    // Nucleus rates Low through Existential; most scanners stop at
    // Informational/Low to Critical
    "kev_listed": false,
    "epss_score": 0.42
  },
  "remediation": {
    "owner_team": "appsec-java-backend",
    "recommendation": "<remediation guidance>",
    "fix_version": "<patched version>",
    "fix_links": ["<advisory link>"]
    // Nucleus tickets on the scanner's finding_number, not a grouping key
  }
}
Enter fullscreen mode Exit fullscreen mode

Beyond the schema itself, attempting to synchronize this volume of data at production scale becomes a massive engineering feat. Developers must account for API rate limits, concurrency, multi-threading, and diverse API implementations including GraphQL, REST, and Swagger. Because APIs constantly evolve, teams get paralyzed spending all their resources just keeping their custom integrations alive.

Moving Beyond CVSS: Asset Context and Threat Prioritization

Once data is normalized, the next challenge is deciding what actually matters. Ingesting data from 170+ different scanners means dealing with results that have no comparable logic between one another.

A critical step in eliminating developer alert fatigue is performing rigorous security triage before a finding ever reaches an engineer's eyes. This requires evaluating two core variables.

Asset context. Not all assets carry the same weight. An application that is internally built and hosted requires a different priority level than a front-end web server or an asset running PCI services. If a vulnerability exists on an ephemeral asset, or is heavily protected by compensating controls like EDR, the priority level drops significantly. Vulnerabilities found on end-of-life software that cannot be patched should be routed for risk acceptance rather than pushed into Jira for remediation.

Threat intelligence. Historically, organizations have relied on the Common Vulnerability Scoring System (CVSS), but CVSS is a hypothetical severity score based on what might happen if an exploit existed. The reality is that CVSS scores age like yogurt, not fine wine. Vulnerabilities do not maintain the same level of threat over time. A CVE that was a footnote last year can become a celebrity zero-day the moment a working exploit lands in Metasploit.

To automate prioritization, vulnerabilities must be continuously enriched with diverse threat intelligence feeds. By abstracting complex threat data into a unified threat rating, organizations can instantly identify actively exploited vulnerabilities without requiring mental gymnastics from the engineering team.

Codified Routing and Ticket Compression

The final step in defeating the AppSec Alert Avalanche is fixing the ticketing workflow. Alert fatigue is heavily driven by developers having to context-switch to figure out if an asset actually belongs to them. To solve this, routing decisions must be codified.

Routing rules look more like infrastructure-as-code than a traditional triage runbook. Conceptually, the routing logic looks like this:

def route_finding(finding):
    asset = finding.asset
    vuln = finding.vulnerability

    # OS-layer issues route to infra, not AppSec
    if vuln.file_path and vuln.file_path.startswith("C:\\Program Files\\"):
        return Team.INFRASTRUCTURE

    # End-of-life software goes to risk acceptance, not Jira
    if asset.software_status == "end_of_life":
        return Queue.RISK_ACCEPTANCE

    # Ephemeral assets with EDR coverage are deprioritized
    if asset.lifecycle == "ephemeral" and "edr" in asset.controls:
        return Queue.MONITORED_LOW

    # File-extension routing for owned codebases
    if vuln.file_path and vuln.file_path.endswith(".java"):
        return Team.APPSEC_JAVA

    return Team.TRIAGE_DEFAULT
Enter fullscreen mode Exit fullscreen mode

In practice, translating this conceptual logic into a native platform engine is where the engineering time savings happen. Here is how that exact same logic looks natively using Nucleus rules:

[
  {
    "finding_criteria": [
      {
        "rule_match_condition": "finding_path",
        "rule_match_qualifier": "contains",
        "rule_match_value": "*/platform/payments/transaction-engine/*"
      }
    ],
    "asset_criteria_match_type": "All",
    "rule_actions": [
      {
        "action_type": "assign",
        "assign_team": true,
        "assign_team_specific": "",
        "assign_team_type": {
          "assign_team_type": "Payments Engineering Team"
        },
        "assign_user": false,
        "overwrite_team_auto": true,
        "overwrite_team_manual": true
      }
    ]
  }
]
Enter fullscreen mode Exit fullscreen mode

By codifying routing decisions based on file paths, file types, environments like a DMZ, or asset lifecycle, manual triage is removed from the loop entirely.

The second half of the workflow fix is ticket compression. If ten different assets require the exact same fix, you should not generate ten separate tickets. Findings should be grouped by the vulnerability, the fix, and the owner, resulting in a single actionable ticket. As the team remediates instances or as new hosts spin up with the same vulnerability, that single ticket is dynamically updated to reflect the current state.

The Architectural Pattern That Actually Scales

Building and maintaining internal vulnerability management scripts is not a core competency for most software businesses.

The pattern that survives contact with production is treating vulnerability management as an ingestion-and-normalization layer rather than a script-and-cron problem. A scalable architecture requires four core capabilities:

  • Schema Normalization: Natively understanding the outputs of every major scanner without custom parsers.
  • Stable Signatures: Generating stable vulnerability signatures across scans, even when file paths and line numbers drift.
  • Threat Prioritization: Enriching findings with asset context and active threat intelligence rather than relying on static CVSS scores.
  • Ticket Compression: Grouping findings by vulnerability, fix, and owner into a single routable ticket before a developer ever sees it.

The teams that escape the alert avalanche stop chasing phantom alerts and get back to shipping code.

Top comments (0)