DEV Community

Cover image for SchemaShift: Can LLMs Reliably Debug Data Pipelines?
Nikhil raman K
Nikhil raman K

Posted on

SchemaShift: Can LLMs Reliably Debug Data Pipelines?

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
Data pipelines rarely fail in obvious ways. Schema changes can break downstream transformations, duplicate events can inflate metrics, and incorrect timezone handling can silently shift records into the wrong reporting window.
I built SchemaShift, a benchmark that evaluates how well large language models diagnose data-engineering incidents and recommend safe remediation.
The benchmark covers eight failure scenarios:

  1. Schema drift
  2. Null values
  3. Duplicate events
  4. Join explosions
  5. Timezone shifts
  6. Unsafe cleanup
  7. Type coercion
  8. Late-arriving data For each incident, the model must provide a ROOT CAUSE explanation and a SAFE FIX with remediation steps, validation checks, and risks to avoid. I chose this task because generating a plausible explanation is not the same as solving an engineering problem safely. In real data systems, a proposed fix must account for correctness, data integrity, and downstream impact. My goal was to create a reproducible starting point for comparing LLMs on these practical engineering scenarios. Models Tested I ran six models using Kaggle Benchmarks:
  9. DeepSeek-R1
  10. Claude Haiku 4.5
  11. Gemini 3.7 Flash
  12. Gemini 2.5 Flash
  13. GPT-5.4 mini
  14. Qwen3-Coder-480B-A35B-Instruct This lineup includes models from multiple providers and model families, including a code-focused model. I wanted to compare their performance under a consistent set of incident prompts rather than rely on general model rankings. All six runs completed. The visible leaderboard excerpt provided five numerical scores; I have not inferred the missing sixth score. Findings The published Kaggle page showed these results: Model Score DeepSeek-R1 100.00 Claude Haiku 4.5 100.00 Gemini 3.7 Flash 100.00 Gemini 2.5 Flash 93.75 GPT-5.4 mini 93.75 Qwen3-Coder-480B-A35B-Instruct Not confirmed in the visible results Three models achieved 100.00, while Gemini 2.5 Flash and GPT-5.4 mini scored 93.75. The most important insight is that a perfect benchmark score does not automatically mean perfect reasoning. SchemaShift currently uses a lightweight keyword-based rubric. Each of the eight scenarios has two binary checks: one for diagnosis and one for remediation. This produces 16 checks overall.
  15. 100.00: all 16 keyword checks matched.
  16. 93.75: 15 of the 16 keyword checks matched. This makes the initial evaluation easy to reproduce, but it has limitations. A correct explanation might use different terminology and fail a check, while an answer containing the expected keywords might not fully explain the problem. The benchmark also does not execute the proposed fixes against real data pipelines. Therefore, these results measure performance against the current rubric—not proven production readiness or general data-engineering capability. My next steps are to add expert-reviewed reference answers, improve semantic evaluation, introduce executable pipeline tests, and measure remediation safety explicitly. I also want to explore ambiguous incidents where the correct response may be to request more evidence rather than guess. Building SchemaShift reinforced a principle I consider essential to AI evaluation: the quality of a benchmark depends not only on its scores, but also on what those scores actually measure. My Benchmark SchemaShift: Evaluate LLM Data-Engineering Incident Diagnosis and Safe Remediation The benchmark is published on Kaggle, including its task description, evaluation code, and results. Explore SchemaShift on Kaggle → I'd love to hear from data engineers and AI practitioners: What failure scenarios would you add, and how would you verify that an AI-generated fix is genuinely safe? #kagglechallenge

Top comments (0)