How a simple entropy-based malware detector evolved into a multi-signal analysis workflow
Introduction
While learning about malware analysis workflows, I kept coming back to the same question:
Do we really need to send every uploaded file into a sandbox before deciding whether it's worth investigating?
Dynamic analysis provides valuable visibility into process creation, network communication, persistence mechanisms, registry modifications, and other behaviors that static analysis cannot easily observe.
The downside is that dynamic analysis is expensive. Running every file through a sandbox consumes time, infrastructure, and analyst attention.
To explore whether some triage decisions could be made earlier, I built a proof-of-concept malware pre-triage pipeline that combines:
- TrID for file identification
- Shannon Entropy for randomness analysis
- Mandiant Capa for capability extraction
The analysis components were validated locally, while the overall architecture was designed around an event-driven AWS deployment model.
This article covers the design decisions, mistakes I made, lessons learned, and how the workflow evolved during testing.
The Initial Assumption
My first implementation was deliberately simple.
The logic looked something like this:
if entropy > 7.2:
quarantine()
At first, it seemed reasonable.
Packed, encrypted, and obfuscated malware often exhibits elevated entropy because the byte stream appears more random.
On paper, a high entropy threshold looked like a quick way to identify suspicious files.
Then I started testing.
The Problem With Entropy
The first thing I discovered was that entropy alone generates a lot of noise.
Many completely legitimate files produced entropy values above my threshold:
- ZIP archives
- Software installers
- Compressed documents
- Packaged applications
For example:
| File Type | Entropy Trend |
|---|---|
| Plain Text | Low |
| ZIP Archive | High |
| Packed Executable | High |
| Encrypted File | High |
The problem became obvious.
A compressed archive and an encrypted malware payload can produce very similar entropy values despite representing completely different security risks.
At that point, entropy stopped being useful as a standalone verdict mechanism.
The challenge shifted from threat detection to false-positive reduction.
Understanding Shannon Entropy
Shannon Entropy measures statistical randomness within a dataset.
The formula is:
Entropy = -Σ P(x) log₂ P(x)
Where:
- P(x) represents the probability of a specific byte value occurring within the file.
Entropy values generally range between:
| Value | Meaning |
|---|---|
| 0.0 | Highly predictable data |
| 8.0 | Highly random data |
In malware analysis, elevated entropy may indicate:
- Packing
- Compression
- Encryption
- Obfuscation
The important word is > may.
Entropy describes the structure of data, not whether that data is malicious.
That distinction became one of the most important lessons of the project.
Why I Added TrID
Once entropy proved insufficient on its own, I needed more context.
The first question became:
What is this file actually supposed to be?
Initially, I considered relying on file extensions.
That quickly seemed unreliable.
A file named:
invoice.pdf
might actually contain:
Win32 Portable Executable
if intentionally disguised.
To address this, I integrated TrID.
TrID identifies files using internal signatures rather than trusting filenames.
This provided several benefits:
- Accurate file identification
- Detection of extension spoofing
- Better interpretation of entropy results
- Additional context for analysts
For example:
Filename: invoice.pdf
TrID Result:
85.4% (.EXE) Win32 Executable
That immediately raises questions worth investigating.
Introducing Capa
At this point, I could identify file types and measure randomness.
However, another important question remained:
What can this executable actually do?
To answer that, I integrated Mandiant Capa.
Capa analyzes executable code and identifies capabilities associated with known behaviors.
Examples include:
- Credential access
- Registry modification
- Persistence mechanisms
- Defense evasion
- Anti-analysis techniques
Rather than looking only at file structure, Capa provides behavioral context.
A file exhibiting elevated entropy and multiple suspicious capabilities is generally more interesting than a file exhibiting elevated entropy alone.
The Workflow That Emerged
By the end of local testing, the workflow looked very different from the original entropy-only design.
The analysis process became:
File
│
▼
TrID File Identification
│
▼
Shannon Entropy Analysis
│
▼
Capa Capability Extraction
│
▼
Risk Evaluation
Instead of relying on a single indicator, the workflow combines multiple signals:
Signal 1: File Identity
What is the file actually?
Signal 2: Entropy
Does the file exhibit characteristics associated with compression, packing, encryption, or obfuscation?
Signal 3: Capability Matches
Does the executable contain capabilities that warrant additional investigation?
No single signal is treated as a definitive answer.
Each contributes context.
Current Implementation
The current proof-of-concept uses:
- Shannon Entropy calculations
- TrID file identification
- Capa capability extraction
The local decision logic currently relies primarily on:
- Entropy threshold evaluation
- Capa rule-match counts
TrID currently provides identification and analyst context, while future iterations could incorporate file-type-aware decision making directly into the risk evaluation process.
Why AWS Lambda?
Although testing was performed locally, the architecture was designed around AWS Lambda from the beginning.
File uploads are naturally event-driven.
When a new object arrives:
- Upload triggers an event.
- Analysis begins automatically.
- Results are generated.
- Suspicious files can be routed for deeper investigation.
Potential advantages include:
- Event-driven execution
- Automatic scaling
- Reduced infrastructure management
- Pay-per-use economics
For lightweight static triage, serverless architecture aligns well with the workload.
- Operational Considerations
Several practical considerations emerged during the design process.
- Dependency Packaging
Both TrID and Capa require additional binaries and supporting resources.
Packaging those dependencies for deployment requires planning.
- Cold Starts
Serverless environments introduce initialization overhead when loading analysis tools and rule sets.
- File Size Constraints
Large archives and software packages may require alternative processing strategies.
- False Positives
False positives remain unavoidable when working with static indicators.
The goal is not perfect detection.
The goal is prioritization.
Future Improvements
If I continue developing this project, I would like to explore:
- Automated archive extraction
- Digital signature validation
- Threat intelligence enrichment
- Risk scoring models
- SIEM integration
- SOAR workflows
Rather than generating simple "clean" or "suspicious" outcomes, future versions could produce weighted risk scores derived from multiple signals.
Key Takeaways
The biggest lesson from this project was that security decisions become more reliable when multiple weak signals are combined.
I learned that:
- Entropy alone generates excessive false positives.
- File identification alone cannot establish intent.
- Capability extraction alone lacks sufficient context.
Combining all three creates a more useful first-pass assessment workflow.
Although this remains a locally validated proof-of-concept, the underlying principle is broadly applicable:
Establish context early, prioritize intelligently, and reserve expensive investigative resources for the cases that genuinely require deeper analysis.
Source Code
GitHub Repository:(https://github.com/Anushkairis/serverless-malware-triage-aws)

Top comments (0)