This is a solo submission by LOI CHIANG HAO for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content
1. What I Built
I built ClashGuard — an autonomous cloud security and compliance conflict arbitration agent powered by Sanity Content Lake and GROQ graph dereferencing (->).
Instead of treating security standards as flat text chunks for a probabilistic search engine, ClashGuard treats competing regulatory authorities, infrastructure targets, and conflicting compliance clauses as first-class relational nodes in a graph.
┌─────────────────────────┐
│ Terraform / K8s Code │
└────────────┬────────────┘
│
▼
┌─────────────────────────┐
│ Static Pattern Traversal│
└────────────┬────────────┘
│ Resource Kind Query
▼
┌─────────────────────────────────────────────────────┐
│ Sanity Content Lake (GROQ) │
│ *[_type=="complianceRule" && defined(conflictsWith)]│
└──────────────────────────┬──────────────────────────┘
│
┌──────────────────┴──────────────────┐
▼ ▼
┌───────────────────────────┐ ┌───────────────────────────┐
│ Authority Source A │ │ Authority Source B │
│ [AWS-WELL-ARCH / TF-REG] │ │ [CIS-AWS / PCI-DSS] │
│ STANCE: RECOMMENDED │ │ STANCE: STRICT_FORBIDDEN │
└─────────────┬─────────────┘ └─────────────┬─────────────┘
│ │
└──────────────────┬──────────────────┘
│
▼
┌─────────────────────────┐
│ Side-by-Side Dual Audit │
│ - Verbatim Evidence │
│ - Multi-Hop Arbitration │
│ - Downloadable SOC Cert │
└─────────────────────────┘
2. The Architectural Crisis: When Authoritative Rulebooks Contradict Each Other
In production cloud engineering, DevOps architects and security compliance officers live in an endless state of friction. The official vendor documentation recommends one configuration pattern for operational scalability, while regulatory frameworks flag that exact pattern as a Level-1 Critical vulnerability.
Consider three classic, multi-million-dollar production clashes:
Clash A: S3 Bucket Public Access vs. CIS AWS Benchmark 2.1.5
-
The Vendor / Architectural Claim:
According to the AWS Well-Architected Framework and standard Terraform CloudFront distribution modules, serving static web assets or media files from an S3 bucket configured with an
s3:GetObjectpublic bucket policy (Principal = "*") is the established, performant pattern. -
The Regulatory Counter-Claim:
The CIS AWS Foundations Benchmark v3.0.0 (Section 2.1.5) strictly mandates that all S3 buckets must enforce
BlockPublicAcls,IgnorePublicAcls,BlockPublicPolicy, andRestrictPublicBuckets. Zero exceptions are granted for CDN origins. Violating this triggers an automated CRITICAL incident across AWS Security Hub. - The Engineering Dilemma: Who is right? If you follow the vendor, you fail the compliance audit. If you follow CIS blindly, you break CDN media distribution.
Clash B: Kubernetes CNI DaemonSets vs. CIS Kubernetes Benchmark 5.2.4
-
The Operational Claim:
Core networking infrastructure DaemonSets (Cilium, Calico, kube-proxy) must run with
hostNetwork: true. Without sharing the host network namespace, eBPF programs cannot bind to the physical host's network interfaces (eth0) to route cluster traffic. -
The Regulatory Counter-Claim:
The CIS Kubernetes Benchmark v1.8.0 (Section 5.2.4) categorically rates
hostNetwork: trueas a HIGH risk container breakout vector. An attacker who compromises a container sharing the host network namespace can inspect host loopback traffic and attack local kubelet services. - The Engineering Dilemma: You cannot run a Kubernetes cluster without a CNI plugin, but standard linters flag the pod manifest as a critical security risk.
Clash C: Remote Analytics Database vs. PCI-DSS 4.0 Requirement 1.3.1
-
The Developer Claim:
The official
terraform-aws-modules/rdsspecification allowspublicly_accessible = truewhen paired with a security group restricting ingress to a trusted external business intelligence CIDR (e.g., Tableau, Looker). - The Regulatory Counter-Claim: PCI-DSS v4.0 (Requirement 1.3.1) categorically prohibits any direct public internet inbound or outbound connections to database components storing cardholder data. Security groups do not satisfy the standard; physical VPC isolation within private subnets is mandatory.
3. Why Traditional Vector Search & Flat RAG Break Down in Production
Most modern "AI compliance assistants" rely on standard Retrieval-Augmented Generation (RAG): chunking documents, embedding them into vector databases via Cosine Similarity, and passing top-K matches into an LLM context window.
This architecture fundamentally collapses when handling contradictory rules:
- Vector Embeddings Have No Directional Relational Semantics: A vector search calculates that the string "Permit GetObject for CloudFront" and "Ensure S3 buckets ignore public ACLs" are semantically close because they both deal with S3 bucket public access policies. However, cosine similarity cannot distinguish between prescriptive permission and strict regulatory prohibition.
-
The LLM Hallucinates Compromises:
When an LLM receives two conflicting, highly authoritative passages in its context window without explicit relational directionality, it hallucinates. It either:
- Falsely assures the user: "Public read is fine because AWS recommends it," creating a severe compliance violation, or
- Falsely panics: "You cannot use CloudFront with S3," paralyzing developer velocity.
The Graph Breakthrough
If keyword or vector search could solve this, we wouldn't need Sanity.
ClashGuard succeeds because the relationship between competing rules is codified as a directed edge in a graph (conflictsWith). The agent does not need to guess whether two rules contradict each other: Sanity's Content Lake already knows.
4. Demo & Interactive Terminal
- 🐙 GitHub Repository: https://github.com/Hao610/clashguard-sanity
- 🗄️ Public Sanity Project ID:
1rqqy3g5 - ⚡ Dataset:
production(Publicly queryable)
Real-Time Code Audit & Side-by-Side Contradiction Panel
ClashGuard features an interactive code audit terminal where developers paste raw Terraform HCL or Kubernetes YAML:
# Example input evaluated by ClashGuard
resource "aws_s3_bucket_public_access_block" "public_cdn" {
bucket = aws_s3_bucket.media_assets.id
block_public_acls = false # Required for legacy CloudFront origin
block_public_policy = false
ignore_public_acls = false
restrict_public_buckets = false
}
The agent scans the AST patterns, executes a multi-hop GROQ graph traversal on Sanity, and outputs:
- Source A (Vendor Guidance): AWS Well-Architected claim + verbatim quotes.
- Source B (Regulatory Audit): CIS Benchmark claim + severity rating.
-
Graph Arbitration Advice: Architectural mitigation paths (e.g., migrating from Principal
*to CloudFront Origin Access Control (OAC)). - Downloadable SOC Compliance Certificate: A verifiable evidence artifact containing transaction hashes and authority links.

Figure 1: ClashGuard's Side-by-Side Contradiction Panel surfacing conflicting claims between AWS Well-Architected and CIS Benchmark, powered by Sanity GROQ graph dereferencing.
5. Code: Traversing the Sanity Graph via GROQ
ClashGuard utilizes a single, elegant GROQ query that traverses from the audited cloud resource to the primary rule, expands its issuing authority, and dereferences all opposing conflict nodes:
*[_type == "complianceRule" && defined(conflictsWith)] {
ruleId,
title,
verdictStance,
verbatimRequirement,
enforcementLevel,
reconciliationAdvice,
"resource": resource-> {
name,
resourceKind,
provider,
riskSurface
},
"source": source-> {
name,
shortCode,
authorityType,
strictnessScore,
homepage,
coreTenet
},
"conflicts": conflictsWith[]-> {
ruleId,
title,
verdictStance,
verbatimRequirement,
enforcementLevel,
reconciliationAdvice,
"source": source-> {
name,
shortCode,
authorityType,
strictnessScore,
homepage,
coreTenet
}
}
}
Why Graph Dereferencing (->) Matters
In standard relational databases or document stores, constructing this query requires multiple SQL JOINs or nested application-layer network requests. With Sanity's GROQ engine, the entire multi-hop conflict topology is dereferenced server-side in a single fast API call.
6. How I Used Sanity
I modeled the compliance domain as an interconnected graph consisting of three distinct document schemas:
Schema 1: authoritySource
Tracks the issuing institution and its governance tier:
-
name: Full title (e.g., CIS AWS Foundations Benchmark v3.0.0). -
authorityType:REGULATORY_COMPLIANCE,VENDOR_RECOMMENDED, orCOMMUNITY_LINTER. -
strictnessScore: Numerical weight (0–100) used by the agent to calibrate risk. -
coreTenet: High-level philosophy (e.g., "Absolute perimeter minimization" vs. "Global media distribution").
Schema 2: cloudResource
Defines the infrastructure boundary being evaluated:
-
name: Target name (AWS S3 Storage Bucket). -
resourceKind: Technical identifier (aws_s3_bucket,kubernetes_pod). -
riskSurface: Attack vectors associated with misconfiguration.
Schema 3: complianceRule (The Relational Core)
Represents the atomic compliance assertion:
-
ruleId: Standard identifier (CIS-AWS-2.1.5). -
verbatimRequirement: The exact, unalterable text of the specification. -
verdictStance:STRICT_FORBIDDENvs.CONDITIONALLY_ALLOWED. -
resource: Reference (->) tocloudResource. -
source: Reference (->) toauthoritySource. -
conflictsWith: Array of references ([]->) pointing directly to contradictory rules in the Content Lake. -
reconciliationAdvice: Grounded architectural recommendations to resolve the clash.
7. Sanity Project Details
As specified by the official challenge requirements:
-
Sanity Project ID:
1rqqy3g5 -
Dataset:
production - Public Read Access: Enabled (Public Dataset)
- Live GROQ Query Endpoint:
https://1rqqy3g5.api.sanity.io/v2024-10-01/data/query/production?query=*[_type=="complianceRule"]
Judges and evaluators can verify the live schema, reference topology, and seeded document mutations directly through the public Sanity API URL above.
8. Why Does Structured Content Matter for AI?
In conversational entertainment or creative writing, an AI hallucination is merely amusing. In cloud infrastructure, regulatory compliance, and cybersecurity, an AI hallucination is an existential risk.
When developers ask an AI to interpret ambiguous security rules:
- Guessing creates catastrophic data breaches.
- Guessing creates multi-million dollar audit fines.
Sanity provides the missing foundation for enterprise AI: Structured Content as Ground Truth. By turning conflicting human policies into an explicit graph of referenced documents, we enable AI agents that refuse to hallucinate, surface both sides of an argument with verifiable citations, and guide developers toward compliant architectures.
Solo submission built by LOI CHIANG HAO for The Sanity Challenge (DEV.to 2026).
Top comments (0)