This is a submission for the Kaggle Benchmarking Challenge
What happens when an LLM recognizes every lexical and semantic signature associated with a vulnerability - IDOR, BOLA, authorization bypass, predictable identifiers - but the available evidence does not actually establish that the vulnerability exists?
That question is the foundation of ProofSec, an evidence-centric security reasoning benchmark designed to evaluate whether LLMs can distinguish security indicators from substantiating evidence, reason under incomplete and contradictory observations, resist terminology and authority bias, incorporate falsifying evidence, and explicitly recognize when a security conclusion is not epistemically justified.
The Problem: Security Reasoning Is Not Pattern Matching
Large language models have become exceptionally capable at semantic retrieval.
Give a model:
GET /api/users/2841/invoices/9281
and introduce:
predictable numeric identifiers
and the model can immediately activate a large cybersecurity concept space:
IDOR
BOLA
Broken Access Control
Authorization Bypass
But there is a fundamental distinction between recognizing the semantic signature of a vulnerability and demonstrating that the underlying security property has actually been violated.
That distinction is where many security reasoning systems become unreliable.
Consider:
Authenticated user: 2841
GET /api/invoices/9281
HTTP/1.1 200 OK
{
"invoice_id": 9281,
"amount": 45000,
"status": "paid"
}
Is this an IDOR?
The evidence is insufficient to establish that conclusion.
The observation does not independently establish:
- who owns invoice
9281 - whether the authenticated principal is authorized to access it
- whether the object belongs to a different principal
- whether an authorization invariant has been violated
- whether access control is enforced upstream or downstream
- whether the endpoint represents the protected resource
- whether the response originated from production data
- whether the response is synthetic, cached, fixture-generated, or otherwise non-authoritative
A predictable identifier is an indicator.
An HTTP 200 response is an observation.
Neither fact, in isolation, constitutes proof of unauthorized cross-principal access.
The central principle behind ProofSec therefore became:
A security indicator is not equivalent to security evidence.
This sounds straightforward.
In practice, it is a surprisingly difficult property for language models to maintain under adversarial framing, incomplete evidence, contradictory observations, and highly salient security terminology.
What I Benchmarked
ProofSec evaluates evidence-sensitive vulnerability classification.
Every scenario requires the model to classify the security state into exactly one canonical category:
Vulnerable
Not Vulnerable
Insufficient Evidence
The third state is deliberately first-class.
Traditional binary vulnerability classification implicitly encourages a forced decision:
YES
or
NO
Real security investigations rarely provide that luxury.
Evidence can be incomplete.
Observations can be ambiguous.
Telemetry can be contradictory.
Authorization boundaries can be unknown.
A response can be suspicious without being conclusive.
ProofSec therefore evaluates whether an LLM can recognize an epistemic boundary:
The available evidence is insufficient to establish the claim.
This changes the benchmark from a conventional vulnerability-recognition exercise into an evaluation of:
- evidentiary sufficiency
- causal relevance
- authorization reasoning
- uncertainty management
- contradiction resolution
- negative evidence integration
- adversarial robustness
- terminology robustness
- controlled hypothesis revision
- evidence-grounded classification
The benchmark is therefore not primarily asking:
"Does the model know what IDOR means?"
It is asking:
"Can the model determine whether the evidence presented actually establishes the security property in question?"
Benchmark Architecture
ProofSec v0.2 contains 110 security reasoning cases.
The corpus is deliberately constructed as an experimental evaluation set rather than an undifferentiated collection of cybersecurity questions.
The benchmark incorporates controlled evaluation families covering:
- One-Fact Flip reasoning
- Evidence ladders
- Contradictory evidence
- Negative evidence
- Authority and terminology robustness
- Baseline security reasoning
The architecture is designed around a core principle:
If the evidentiary state changes, the model should be sensitive to that change.
Conceptually:
ProofSec
│
┌──────────────┼──────────────┐
│ │ │
One-Fact Flip Evidence State Contradiction
│ │ │
└──────────────┼──────────────┘
│
Security Scenario
│
▼
LLM Inference
│
▼
Structured Assessment
│
┌──────────────┼──────────────┐
│ │ │
Classification Evidence State Explanation
│
▼
Deterministic Assertion
│
▼
Numerical Score
The implementation deliberately isolates:
Dataset
↓
Prompt Construction
↓
Model Inference
↓
Structured Parsing
↓
Assertion
↓
Scoring
This separation is not cosmetic.
It establishes distinct failure domains.
A dataset defect is not a model failure.
A packaging failure is not a reasoning failure.
A schema violation is not necessarily a classification failure.
A scoring-interface defect is not a model-performance measurement.
That distinction became one of the most important engineering principles in the project.
1. One-Fact Perturbation Testing
One of the core mechanisms in ProofSec is controlled one-fact perturbation.
Instead of constructing completely unrelated questions, ProofSec creates scenarios where a security-relevant fact changes while much of the surrounding semantic structure remains invariant.
For example:
Case A
Authenticated user = 2841
Invoice owner = 2841
Authorization rule:
invoice.owner_id == authenticated_user.id
Case B
Authenticated user = 2841
Invoice owner = 9127
Authorization rule:
invoice.owner_id == authenticated_user.id
The semantic surface can remain substantially similar.
The critical security relationship changes.
That gives us a controlled counterfactual:
Δ Input
↓
Δ Security-Relevant Fact
↓
Expected Δ Security State
This is fundamentally more informative than asking:
"Is this an IDOR?"
The actual evaluation question becomes:
Did the model detect the fact that causally changes the security state?
This introduces a form of counterfactual sensitivity testing.
If one authorization-relevant fact changes and the model preserves the same classification, the benchmark exposes a specific failure mode.
The model may understand the terminology.
It may understand the vulnerability class.
But it may not be correctly conditioning its conclusion on the evidence that actually determines the security property.
2. Evidence-State Modeling
ProofSec explicitly models the evidentiary state of a scenario.
The benchmark uses states including:
WEAK
PARTIAL
DECISIVE
CONTRADICTORY
NEGATIVE
UNKNOWN
This creates an evidence hierarchy rather than treating every security-relevant observation as equally probative.
A simplified progression is:
Predictable identifier
│
▼
Weak security indicator
│
▼
Ownership relationship established
│
▼
Cross-principal access demonstrated
│
▼
Authorization invariant violated
│
▼
DECISIVE evidence
Consider the difference between:
Predictable ID
and:
Unauthorized cross-user object access
These observations possess radically different evidentiary weight.
The first may justify investigation.
The second can establish a concrete violation of an authorization property.
ProofSec therefore attempts to prevent a model from collapsing the following distinction:
Suspicion
≠
Evidence
≠
Proof
That distinction is foundational to trustworthy security analysis.
3. Contradictory Evidence
Real security investigations are not monotonic.
Evidence can conflict.
Telemetry can be stale.
Caches can contain misleading artifacts.
Fixtures can resemble production responses.
A preliminary observation can subsequently be invalidated by a higher-authority observation.
ProofSec therefore incorporates contradiction-resolution scenarios.
A simplified reasoning sequence might look like:
Initial observation
│
▼
HTTP 200 response
│
▼
Potential authorization anomaly
│
├───────────────┐
│ │
▼ ▼
Synthetic fixture Actual endpoint
│ │
▼ ▼
Cached response HTTP 403
│ │
└───────┬───────┘
▼
Re-evaluate hypothesis
The important capability is not simply detecting the first suspicious observation.
The model must determine whether later evidence changes the validity of the original hypothesis.
This introduces a critical distinction between:
Evidence accumulation
and:
Evidence revision
A model that merely accumulates confirming signals can behave very differently from one capable of hypothesis revision under contradictory evidence.
ProofSec deliberately tests that boundary.
4. Negative Evidence
Many vulnerability benchmarks are heavily oriented toward positive findings.
ProofSec deliberately gives negative evidence first-class status.
For example:
User A requests User B's object
│
▼
Authorization rule verified
│
▼
Cross-user request returns 403
│
▼
Unauthorized access not demonstrated
A security reasoning system must be able to update its hypothesis in both directions.
It must identify:
Evidence supporting vulnerability
and:
Evidence weakening or falsifying vulnerability hypothesis
This matters because security investigation is fundamentally adversarial.
The objective is not to collect evidence that confirms the initial hypothesis.
The objective is to determine whether the hypothesis survives attempts to falsify it.
That makes negative evidence an important component of hypothesis discrimination.
5. Authority Bias and Terminology Robustness
Another failure mode targeted by ProofSec is authority and terminology bias.
Cybersecurity vocabulary carries enormous semantic weight.
Terms such as:
critical
confirmed
exploit
CVE
researcher
privilege escalation
authorization bypass
security issue
can exert disproportionate influence over LLM outputs.
ProofSec therefore evaluates whether the model remains anchored to technical evidence when the surrounding terminology changes.
The underlying principle is:
Terminology describes a claim. Evidence substantiates it.
Consider two descriptions:
Possible authorization issue
and:
Confirmed critical authorization vulnerability
If the underlying technical evidence is unchanged, the semantic framing should not arbitrarily alter the actual security state.
This probes linguistic framing sensitivity and authority-induced classification drift.
The benchmark is therefore not only testing cybersecurity knowledge.
It is testing whether that knowledge remains subordinate to the evidentiary record.
Structured Outputs
I intentionally avoided unconstrained free-form text as the primary evaluation interface.
Each model response is constrained into a structured schema containing:
classification
evidence_state
supporting_evidence
missing_evidence
safe_verification
impact
This creates a deterministic machine-readable boundary between inference and evaluation.
The primary scoring target is:
classification
The remaining fields provide diagnostic observability.
For example:
classification:
Vulnerable
could be accompanied by:
supporting_evidence:
"The identifier is sequential."
That exposes a potentially significant evidentiary defect:
Predictable Identifier
≠
Proven Unauthorized Access
The final classification alone cannot reveal whether the model arrived at the conclusion through valid or invalid reasoning.
Structured assessment therefore provides a richer diagnostic surface.
Deterministic Ground Truth
I deliberately avoided making another LLM the primary arbiter of the classification.
ProofSec uses canonical ground-truth labels and deterministic assertions.
The core evaluation chain is:
Scenario
↓
Model
↓
Structured Classification
↓
Canonical Ground Truth
↓
Exact Assertion
↓
Score
rather than:
Scenario
↓
Model A
↓
Model B judges Model A
↓
Score
For the core classification metric, deterministic evaluation establishes a cleaner experimental boundary.
It also reduces the possibility of evaluation contamination, where the judgment model introduces another layer of model-dependent interpretation into the primary metric.
Dataset Integrity
A benchmark is only scientifically useful if its evaluation corpus is reproducible.
ProofSec v0.2 uses a frozen dataset representation with a SHA-256 integrity digest:
422501a4db424c30c8ef24b61183351ec8a4bd2096e2671cf0e6bdf91e133a80
The project maintains a manifest and an integrity-verification workflow.
The intended reproducibility invariant is:
Same benchmark version
+
Same task corpus
+
Same evaluation protocol
=
Comparable experiment
Without corpus integrity, two apparently identical experiments can silently evaluate different datasets.
That means dataset provenance, versioning, and integrity verification are part of the experimental methodology itself.
The Benchmark Itself Became an Experiment
One of the most valuable discoveries during the project was that benchmark engineering is itself part of evaluation science.
My initial task implementation used:
-> None
with assertions.
That produced pass/fail-style behavior.
I initially expected the benchmark to expose a numerical accuracy metric.
That assumption was incorrect.
I subsequently changed the task interface to:
-> float
and explicitly returned:
accuracy = correct / total
return float(accuracy)
This exposed a broader principle:
The semantics of the evaluation harness are part of the experimental design.
A benchmark can execute successfully while still reporting an invalid measurement if the scoring contract is incorrectly defined.
In other words:
Correct Inference
≠
Correct Evaluation
A trustworthy benchmark requires both.
The Kaggle Runtime Was Another Experimental Variable
One early published iteration failed with:
NameError: name 'ALL_TASKS_JSON' is not defined
This was not a model reasoning failure.
The model never reached the reasoning stage.
The task failed during execution before the first model inference step.
That forced a more rigorous separation of:
Dataset Integrity
│
▼
Task Packaging
│
▼
Runtime Execution
│
▼
Model Inference
│
▼
Structured Parsing
│
▼
Assertion
│
▼
Scoring
This distinction is fundamental.
A benchmark must not silently transform:
Runtime Error
into:
Model Error
Those are different failure classes with different remediation paths.
It also exposed an important deployment boundary.
My local development environment contained the canonical task source under:
C:\Users\User\Documents\ProofSec\kaggle\tasks
but a Kaggle-hosted execution environment cannot directly access that Windows filesystem.
The benchmark therefore has to package the necessary evaluation corpus and task implementation into the executable artifact available to the remote runtime.
That became an important lesson in benchmark portability and execution-environment isolation.
Versioned Benchmark Development
The benchmark went through multiple iterations while I validated the evaluation infrastructure.
One development version displayed assertion counts including:
Claude Sonnet 4.6 0 Pass
GPT-5.5 75 Pass
Gemini 3.5 Flash 95 Pass
Gemini 3.7 Flash 87 Pass
A subsequent controlled version successfully executed its smoke-test assertions across the evaluated models.
These development-stage measurements should not be interpreted as a universal model ranking.
They represent different stages of benchmark engineering.
During this process I was simultaneously validating:
- dataset construction
- task implementation
- structured-output schema
- assertion semantics
- scoring semantics
- remote runtime packaging
- model inference
This separation is essential for responsible interpretation.
A benchmark-development artifact is not automatically equivalent to a final scientific measurement.
Local Validation
Before treating Kaggle execution as authoritative, I also used local evaluation to inspect the complete 110-case corpus.
Representative local validation produced:
| Model | Correct | Total | Accuracy |
|---|---|---|---|
| Claude Sonnet 4.6 | 80 | 110 | 72.73% |
| Gemini 3.7 Flash | 86 | 110 | 78.18% |
| Gemini 3.5 Flash | 84 | 110 | 76.36% |
| GPT-5.5 | 67 | 110 | 60.91% |
These figures are local validation measurements under the specific execution conditions used during development.
They should not be interpreted as immutable measurements of model capability.
LLM inference can vary because of:
- model revisions
- provider routing
- stochastic inference
- runtime configuration
- API behavior
- execution environment
- infrastructure conditions
This variability is itself relevant to reproducibility.
The objective was therefore not to reduce the benchmark to:
"Model X is better."
The objective was to characterize failure modes, evidentiary behavior, and reasoning sensitivity under a controlled corpus.
A Particularly Interesting Failure Mode
One of the recurring failure patterns motivating ProofSec can be represented as:
Security Vocabulary Recognition
↓
Semantic Activation
↓
Premature Classification
↓
Evidence Is Never Actually Tested
The desired reasoning pathway is:
Scenario
↓
Extract Security-Relevant Claims
↓
Identify Actors and Authorization Boundaries
↓
Identify Directly Observed Facts
↓
Separate Indicators from Decisive Evidence
↓
Evaluate Contradictory Evidence
↓
Evaluate Negative Evidence
↓
Determine Evidentiary Sufficiency
↓
Classify
The distinction is fundamental.
A model can possess extensive cybersecurity knowledge while still being unreliable at evidence-grounded adjudication.
This is the capability ProofSec attempts to isolate.
Why "Insufficient Evidence" Matters
This is one of the most consequential design decisions in ProofSec.
A conventional benchmark might ask:
Is the system vulnerable?
YES / NO
ProofSec instead represents:
Vulnerable
Not Vulnerable
Insufficient Evidence
This explicitly models epistemic uncertainty.
The model must distinguish between:
Observed
and:
Established
That distinction matters operationally.
Consider:
Unsupported Finding
↓
Analyst Investigation
↓
Engineering Interruption
↓
Potential Remediation
↓
Operational Cost
A system that generates large quantities of unsupported security findings can impose substantial downstream cost even when its vulnerability-recognition capability appears impressive.
Abstention is therefore not merely a failure to classify.
Under uncertainty, abstention can be a legitimate and necessary security behavior.
The Benchmark Is Essentially Testing Epistemic Discipline
At a deeper level, ProofSec is not simply a cybersecurity benchmark.
It evaluates whether an LLM can maintain an explicit boundary between:
What do I know?
↓
What does the evidence establish?
↓
What remains unknown?
↓
What conclusion is justified?
This is fundamentally an epistemic reasoning problem.
Cybersecurity provides a particularly concrete environment in which to measure it because security conclusions often have direct operational consequences.
The same methodology could potentially extend to:
- code review
- incident response
- fraud detection
- compliance analysis
- threat intelligence
- infrastructure diagnostics
- financial anomaly investigation
- legal document analysis
- scientific reasoning
The security domain is simply where I chose to operationalize the problem.
Engineering Stack
ProofSec is implemented using:
- Python
- Kaggle Benchmarks
kaggle_benchmarks- Pydantic
- structured model outputs
- deterministic assertions
- JSON-based frozen corpus
- SHA-256 integrity verification
- controlled perturbation generation
- evidence-state annotations
- local validation
- Kaggle-hosted execution
The architecture maintains explicit boundaries between:
DATA
PROMPT CONSTRUCTION
MODEL INFERENCE
STRUCTURED PARSING
EVALUATION
SCORING
This makes the system substantially easier to:
- debug
- version
- reproduce
- audit
- extend
- attribute failures within
The benchmark is therefore treated as an engineered evaluation artifact rather than merely a collection of prompts.
Reproducibility and Development Workflow
The development workflow became:
Security Hypothesis
↓
Construct Controlled Scenario
↓
Define Canonical Ground Truth
↓
Generate Adversarial / Paired Variant
↓
Validate Dataset
↓
Freeze Corpus
↓
Generate Integrity Hash
↓
Implement Structured Task
↓
Run Local Validation
↓
Validate Kaggle Packaging
↓
Execute Against Multiple Models
↓
Inspect Assertions
↓
Analyze Failure Modes
↓
Revise Benchmark Infrastructure
This treats the benchmark as a versioned research artifact.
The project structure is maintained as an engineering repository containing the benchmark implementation, evaluation infrastructure, integrity controls, and supporting documentation.
The objective is to make the benchmark auditable rather than opaque.
What I Learned
1. Security reasoning is not vulnerability keyword matching
The presence of:
ID
endpoint
HTTP 200
user
invoice
does not automatically establish IDOR.
The authorization relationship is the security property that matters.
2. Counterfactual sensitivity is extremely valuable
If a single security-relevant fact changes, the model should respond to that change.
Controlled perturbation therefore provides a much stronger probe of reasoning sensitivity than simply increasing the number of unrelated benchmark questions.
3. Negative evidence deserves first-class treatment
A security reasoning system must be capable of recognizing evidence that weakens or falsifies its initial hypothesis.
4. Contradictions expose hypothesis-update behavior
A model that performs well under clean, monotonic evidence may behave very differently when observations conflict.
Contradiction handling is therefore a meaningful robustness dimension.
5. Abstention is measurable
"Insufficient Evidence" can be treated as an explicit epistemic state rather than an evasive response.
6. Structured outputs improve diagnosability
A classification tells us:
What did the model decide?
A structured assessment can additionally expose:
What evidence did it identify?
What evidence did it consider missing?
What verification did it propose?
That creates a substantially richer diagnostic surface.
7. The benchmark harness must be engineered like software
Dataset defects, packaging failures, schema violations, runtime exceptions, parsing failures, and model reasoning failures require separate failure categories.
8. Reproducibility is part of validity
A benchmark without corpus integrity, version control, and execution controls can silently drift.
A score without provenance is considerably less useful than a score attached to a reproducible experimental configuration.
What I Want to Measure Next
ProofSec v0.2 focuses on controlled evidence reasoning.
The next iteration can extend the methodology considerably.
Temporal Reasoning
Instead of presenting all evidence simultaneously:
T0 → Initial Telemetry
T1 → Additional Observation
T2 → Contradictory Artifact
T3 → Authorization Test
I want to evaluate whether the model updates its hypothesis correctly as evidence arrives.
This transforms static classification into a sequential evidence-update problem.
Confidence Calibration
Instead of evaluating only:
classification
evaluate:
classification
+
confidence
Then measure whether confidence is statistically aligned with empirical correctness.
A model that is wrong with extremely high confidence represents a fundamentally different operational risk profile from a model that identifies uncertainty.
Evidence Attribution
Require the model to answer:
Which exact fact changed your classification?
This enables a deeper evaluation:
Correct Conclusion
+
Correct Evidentiary Basis
versus:
Correct Conclusion
+
Incorrect Evidentiary Basis
This distinction matters because a model can occasionally arrive at the correct answer for the wrong reasons.
Cross-Domain Generalization
The framework can be extended into:
SSRF
Authentication Bypass
Privilege Escalation
CSRF
Path Traversal
Race Conditions
Business Logic Flaws
Cloud IAM
API Authorization
Multi-Tenant Isolation
The broader research question becomes:
Is evidence-grounded security reasoning a generalizable capability, or is it primarily a collection of domain-specific semantic associations?
The Bigger Question
When I started building ProofSec, the obvious question was:
Which LLM is better at cybersecurity?
After engineering the benchmark, that no longer seemed like the most interesting question.
The more fundamental question is:
Can we construct evaluations that determine whether an AI security conclusion is actually justified by the evidence available to it?
A leaderboard such as:
Model A - 82%
Model B - 79%
Model C - 76%
is useful.
But it is incomplete.
For an AI security system, I want to know:
What evidence did the model use?
What evidence did it ignore?
Did it infer facts that were never provided?
Did it confuse an indicator with proof?
Did it recognize negative evidence?
Did it resolve contradictions?
Did terminology alter its classification?
Did it respond to a one-fact perturbation?
Did it revise its hypothesis?
Did it recognize when the evidence was insufficient?
Those questions are much closer to the requirements of a real security-analysis system.
My Benchmark
ProofSec - Security Evidence Reasoning Benchmark
ProofSec v0.2 contains 110 security reasoning cases designed around controlled evidence manipulation, structured classification, deterministic ground truth, integrity verification, and reproducible multi-model evaluation.
The benchmark is built around one principle:
Security conclusions should be proportional to the evidence supporting them.
ProofSec operationalizes that principle through:
- controlled one-fact perturbations
- evidence-state modeling
- contradiction resolution
- negative evidence
- authority and terminology robustness
- structured outputs
- deterministic assertions
- frozen corpus integrity
- SHA-256 verification
- cross-model execution
Open-Source Implementation
The benchmark implementation and supporting engineering work are available on GitHub:
The repository is intended to make the benchmark more than a Kaggle submission.
It provides an engineering artifact around the evaluation methodology, allowing the benchmark to be inspected, versioned, reproduced, and extended independently of the Kaggle interface.
The broader objective is to preserve the chain:
Research Hypothesis
↓
Dataset Construction
↓
Ground Truth
↓
Evaluation Harness
↓
Integrity Verification
↓
Model Execution
↓
Results
↓
Failure Analysis
That provenance matters.
A benchmark should not only report a number.
It should make it possible to understand where that number came from.
Final Takeaway
The most important result from building ProofSec was not a single model score.
It was discovering how easily an evaluation can appear to measure security reasoning while actually measuring something else.
A benchmark can accidentally measure:
Keyword Recognition
instead of:
Evidence Reasoning
It can measure:
Runtime Correctness
instead of:
Model Correctness
It can measure:
Confidence
instead of:
Evidence-Justified Certainty
It can measure:
Semantic Familiarity
instead of:
Causal Sensitivity
And it can report a precise-looking number while the underlying dataset, runtime, scoring semantics, or evaluation protocol is not actually controlled.
ProofSec is my attempt to make those boundaries explicit.
The objective is not simply to build an LLM that identifies the largest number of possible vulnerabilities.
The objective is to construct an evaluation framework capable of answering a substantially harder question:
When an AI claims that a security vulnerability exists, did the available evidence actually justify that claim?
That is the standard I want AI-assisted security tooling to eventually be held to.
Not:
Did the model recognize the vulnerability vocabulary?
But:
Did the model establish the security claim from the evidence?
That distinction is the entire premise of ProofSec.
Resources
ProofSec Benchmark
ProofSec Source Repository
Kaggle Benchmarking Challenge
Kaggle Benchmarking Challenge on DEV
Top comments (0)