DEV Community

Dakota Ma profile picture

Dakota Ma

Frontend developer crafting beautiful user experiences.

Location San Francisco, CA Joined Joined on 
A Perfect Pass Rate Is a Blind Harness

A Perfect Pass Rate Is a Blind Harness

Comments
9 min read
Keep Schema Failures Out of Your Quality Score

Keep Schema Failures Out of Your Quality Score

Comments
7 min read
Tuning-Set Pass Rate Is Not an Eval

Tuning-Set Pass Rate Is Not an Eval

Comments
7 min read
The Completion Matched. The Envelope Did Not.

The Completion Matched. The Envelope Did Not.

Comments 1
8 min read
Treat the Grader as Code, Not a Hidden Prompt

Treat the Grader as Code, Not a Hidden Prompt

Comments
6 min read
Calibrate the Noise Floor of Your Eval Harness

Calibrate the Noise Floor of Your Eval Harness

Comments
6 min read
Fail the Eval When Samples Disagree

Fail the Eval When Samples Disagree

Comments
6 min read
Diff Last-Good Completions, Not Just Pass Rate

Diff Last-Good Completions, Not Just Pass Rate

Comments
7 min read
Mutate Golden Cases Before You Trust the Eval

Mutate Golden Cases Before You Trust the Eval

Comments
6 min read
When Context Is Incomplete, Completion Is the Bug

When Context Is Incomplete, Completion Is the Bug

Comments
6 min read
Pin Behaviors Across Model Swaps

Pin Behaviors Across Model Swaps

Comments
8 min read
Version Golden Cases Like Schema Migrations

Version Golden Cases Like Schema Migrations

Comments
5 min read
Put Forbidden Inferences in the Golden File

Put Forbidden Inferences in the Golden File

Comments
6 min read
Score the Trace, Not the Final Payload

Score the Trace, Not the Final Payload

Comments
7 min read
Score the Trace, Not the Last Message

Score the Trace, Not the Last Message

Comments
7 min read
Catch Tool Calls That Invent Missing Arguments

Catch Tool Calls That Invent Missing Arguments

Comments 1
7 min read
Schema-First Validation: Catch Silent Model Drift

Schema-First Validation: Catch Silent Model Drift

Comments
7 min read
Free Tiers Are Just Queues: A Capacity Plan for Token-Limited Pipelines

Free Tiers Are Just Queues: A Capacity Plan for Token-Limited Pipelines

Comments
5 min read
Treat LLM Output Like an API Contract

Treat LLM Output Like an API Contract

Comments
5 min read
Your Paid Model Needs a Cheap Skeptic

Your Paid Model Needs a Cheap Skeptic

Comments
3 min read
Your Grader Is Drifting Too: A Self-Auditing Prompt Eval Harness

Your Grader Is Drifting Too: A Self-Auditing Prompt Eval Harness

Comments
5 min read
A Cache-First LLM Gateway You Can Run on a Free Server

A Cache-First LLM Gateway You Can Run on a Free Server

Comments
5 min read
Golden Sets Are the Unit Tests Your LLM Feature Never Had

Golden Sets Are the Unit Tests Your LLM Feature Never Had

Comments
4 min read
Prompt Drift Is a Quiet Deployment Bug: A Nightly Check That Costs Nothing

Prompt Drift Is a Quiet Deployment Bug: A Nightly Check That Costs Nothing

Comments 1
4 min read
Silent Regressions Have No Stack Trace: A Minimal Prompt Eval Harness

Silent Regressions Have No Stack Trace: A Minimal Prompt Eval Harness

Comments 2
5 min read
The Prompt Changed. Nothing Broke. That's the Problem.

The Prompt Changed. Nothing Broke. That's the Problem.

Comments
5 min read
A Quota-Aware Proxy for Free AI Endpoints: Treating Limits as a Budget, Not a Wall

A Quota-Aware Proxy for Free AI Endpoints: Treating Limits as a Budget, Not a Wall

Comments
3 min read
A Free-Tier AI PR Reviewer: A GitHub Actions Workflow That Actually Works

A Free-Tier AI PR Reviewer: A GitHub Actions Workflow That Actually Works

1
Comments
4 min read
The Model Was Fine. My Token Assumptions Weren't.

The Model Was Fine. My Token Assumptions Weren't.

Comments
5 min read
The Crash That Only Happened in Production

The Crash That Only Happened in Production

Comments
3 min read
I Almost Posted a Hot Take About a Cheap New Model. Then I Built a Free Triage Harness.

I Almost Posted a Hot Take About a Cheap New Model. Then I Built a Free Triage Harness.

Comments
3 min read
Shadow-Test a New Coding Model in One Week: A Correction-Log Method

Shadow-Test a New Coding Model in One Week: A Correction-Log Method

Comments
4 min read
Your Agent Reads Untrusted Text All Day. Here's How I Grade What It Does With It.

Your Agent Reads Untrusted Text All Day. Here's How I Grade What It Does With It.

Comments
6 min read
loading...