DEV Community

Anushkairis
Anushkairis

Posted on

Building a Malware Pre-Triage Pipeline with TrID, Capa, and Shannon Entropy

How a simple entropy-based malware detector evolved into a multi-signal analysis workflow

Introduction

While learning about malware analysis workflows, I kept coming back to the same question:

Do we really need to send every uploaded file into a sandbox before deciding whether it's worth investigating?

Dynamic analysis provides valuable visibility into process creation, network communication, persistence mechanisms, registry modifications, and other behaviors that static analysis cannot easily observe.

The downside is that dynamic analysis is expensive. Running every file through a sandbox consumes time, infrastructure, and analyst attention.

To explore whether some triage decisions could be made earlier, I built a proof-of-concept malware pre-triage pipeline that combines:

  • TrID for file identification
  • Shannon Entropy for randomness analysis
  • Mandiant Capa for capability extraction

The analysis components were validated locally, while the overall architecture was designed around an event-driven AWS deployment model.

This article covers the design decisions, mistakes I made, lessons learned, and how the workflow evolved during testing.

Architecture overview

The Initial Assumption

My first implementation was deliberately simple.

The logic looked something like this:

if entropy > 7.2:
quarantine()

At first, it seemed reasonable.

Packed, encrypted, and obfuscated malware often exhibits elevated entropy because the byte stream appears more random.

On paper, a high entropy threshold looked like a quick way to identify suspicious files.

Then I started testing.

The Problem With Entropy

The first thing I discovered was that entropy alone generates a lot of noise.

Many completely legitimate files produced entropy values above my threshold:

  • ZIP archives
  • Software installers
  • Compressed documents
  • Packaged applications

For example:

File Type Entropy Trend
Plain Text Low
ZIP Archive High
Packed Executable High
Encrypted File High

The problem became obvious.

A compressed archive and an encrypted malware payload can produce very similar entropy values despite representing completely different security risks.

At that point, entropy stopped being useful as a standalone verdict mechanism.

The challenge shifted from threat detection to false-positive reduction.

Understanding Shannon Entropy
Shannon Entropy measures statistical randomness within a dataset.

The formula is:

Entropy = -Σ P(x) log₂ P(x)

Where:

  • P(x) represents the probability of a specific byte value occurring within the file.

Entropy values generally range between:

Value Meaning
0.0 Highly predictable data
8.0 Highly random data

In malware analysis, elevated entropy may indicate:

  • Packing
  • Compression
  • Encryption
  • Obfuscation

The important word is > may.

Entropy describes the structure of data, not whether that data is malicious.

That distinction became one of the most important lessons of the project.

Why I Added TrID

Once entropy proved insufficient on its own, I needed more context.

The first question became:
What is this file actually supposed to be?

Initially, I considered relying on file extensions.

That quickly seemed unreliable.

A file named:

invoice.pdf

might actually contain:
Win32 Portable Executable

if intentionally disguised.

To address this, I integrated TrID.
TrID identifies files using internal signatures rather than trusting filenames.

This provided several benefits:

  • Accurate file identification
  • Detection of extension spoofing
  • Better interpretation of entropy results
  • Additional context for analysts

For example:

Filename: invoice.pdf

TrID Result:
85.4% (.EXE) Win32 Executable

That immediately raises questions worth investigating.

Introducing Capa

At this point, I could identify file types and measure randomness.

However, another important question remained:
What can this executable actually do?

To answer that, I integrated Mandiant Capa.

Capa analyzes executable code and identifies capabilities associated with known behaviors.

Examples include:

  • Credential access
  • Registry modification
  • Persistence mechanisms
  • Defense evasion
  • Anti-analysis techniques

Rather than looking only at file structure, Capa provides behavioral context.

A file exhibiting elevated entropy and multiple suspicious capabilities is generally more interesting than a file exhibiting elevated entropy alone.

The Workflow That Emerged

By the end of local testing, the workflow looked very different from the original entropy-only design.

The analysis process became:

File
│
▼
TrID File Identification
│
▼
Shannon Entropy Analysis
│
▼
Capa Capability Extraction
│
▼
Risk Evaluation

Instead of relying on a single indicator, the workflow combines multiple signals:

Signal 1: File Identity
What is the file actually?

Signal 2: Entropy
Does the file exhibit characteristics associated with compression, packing, encryption, or obfuscation?

Signal 3: Capability Matches
Does the executable contain capabilities that warrant additional investigation?

No single signal is treated as a definitive answer.
Each contributes context.

Current Implementation
The current proof-of-concept uses:

  • Shannon Entropy calculations
  • TrID file identification
  • Capa capability extraction

The local decision logic currently relies primarily on:

  • Entropy threshold evaluation
  • Capa rule-match counts

TrID currently provides identification and analyst context, while future iterations could incorporate file-type-aware decision making directly into the risk evaluation process.

Why AWS Lambda?

Although testing was performed locally, the architecture was designed around AWS Lambda from the beginning.

File uploads are naturally event-driven.

When a new object arrives:

  • Upload triggers an event.
  • Analysis begins automatically.
  • Results are generated.
  • Suspicious files can be routed for deeper investigation.

Potential advantages include:

  • Event-driven execution
  • Automatic scaling
  • Reduced infrastructure management
  • Pay-per-use economics

For lightweight static triage, serverless architecture aligns well with the workload.

- Operational Considerations
Several practical considerations emerged during the design process.

- Dependency Packaging
Both TrID and Capa require additional binaries and supporting resources.
Packaging those dependencies for deployment requires planning.

- Cold Starts
Serverless environments introduce initialization overhead when loading analysis tools and rule sets.

- File Size Constraints
Large archives and software packages may require alternative processing strategies.

- False Positives
False positives remain unavoidable when working with static indicators.

The goal is not perfect detection.
The goal is prioritization.

Future Improvements

If I continue developing this project, I would like to explore:

  • Automated archive extraction
  • Digital signature validation
  • Threat intelligence enrichment
  • Risk scoring models
  • SIEM integration
  • SOAR workflows

Rather than generating simple "clean" or "suspicious" outcomes, future versions could produce weighted risk scores derived from multiple signals.

Key Takeaways

The biggest lesson from this project was that security decisions become more reliable when multiple weak signals are combined.

I learned that:

  • Entropy alone generates excessive false positives.
  • File identification alone cannot establish intent.
  • Capability extraction alone lacks sufficient context.

Combining all three creates a more useful first-pass assessment workflow.

Although this remains a locally validated proof-of-concept, the underlying principle is broadly applicable:

Establish context early, prioritize intelligently, and reserve expensive investigative resources for the cases that genuinely require deeper analysis.

Source Code
GitHub Repository:(https://github.com/Anushkairis/serverless-malware-triage-aws)

Top comments (0)