Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
← All Trends
Custom Local Evals for New AI Coding Models
77 posts in this trend in the last 7 days
•
Active 4 minutes ago
A Reproducible Harness for Evaluating New Open Models Before They Touch Your Codebase
Dakota Lin
Dakota Lin
Dakota Lin
Follow
Aug 10
A Reproducible Harness for Evaluating New Open Models Before They Touch Your Codebase
#
opensource
#
ai
#
testing
#
productivity
Comments
Add Comment
4 min read
A New Cheap Model Dropped This Week. Here's the Harness I Run Before I Switch Anything
Riley Wang
Riley Wang
Riley Wang
Follow
Aug 13
A New Cheap Model Dropped This Week. Here's the Harness I Run Before I Switch Anything
#
ai
#
testing
#
productivity
#
tooling
Comments
Add Comment
5 min read
Before You Switch to the Hottest Open-Weight Model, Run This 30-Minute Eval Harness
Taylor Wang
Taylor Wang
Taylor Wang
Follow
Aug 10
Before You Switch to the Hottest Open-Weight Model, Run This 30-Minute Eval Harness
#
ai
#
opensource
#
programming
#
productivity
Comments
Add Comment
4 min read
A New Open Model Dropped. Here's My 30-Minute Reproducible Eval Before I Trust It With My Codebase
Dakota Liu
Dakota Liu
Dakota Liu
Follow
Aug 10
A New Open Model Dropped. Here's My 30-Minute Reproducible Eval Before I Trust It With My Codebase
#
ai
#
opensource
#
llm
#
productivity
1
reaction
Comments
Add Comment
5 min read
Stop Re-litigating Every Model Launch: A Personal Bake-Off Rig You Can Rerun Forever
Emery Chen
Emery Chen
Emery Chen
Follow
Aug 10
Stop Re-litigating Every Model Launch: A Personal Bake-Off Rig You Can Rerun Forever
#
ai
#
productivity
#
testing
#
programming
Comments
Add Comment
6 min read
Vibes Are Not a Benchmark: A 30-Minute Harness to Test Whether a Free Coding Model Can Touch Your Repo
Dakota Huang
Dakota Huang
Dakota Huang
Follow
Aug 12
Vibes Are Not a Benchmark: A 30-Minute Harness to Test Whether a Free Coding Model Can Touch Your Repo
#
ai
#
productivity
#
python
#
tooling
Comments
2
comments
5 min read
New Open-Weight Model Drops Every Week. Here's a Reproducible Way to Decide If It Belongs in Your Workflow
Finley Zhu
Finley Zhu
Finley Zhu
Follow
Aug 10
New Open-Weight Model Drops Every Week. Here's a Reproducible Way to Decide If It Belongs in Your Workflow
#
ai
#
opensource
#
programming
#
productivity
Comments
Add Comment
5 min read
Don't Trust the Demo: A Repeatable Test Harness for Evaluating Free AI Coding Models
Finley Zhou
Finley Zhou
Finley Zhou
Follow
Aug 10
Don't Trust the Demo: A Repeatable Test Harness for Evaluating Free AI Coding Models
#
ai
#
programming
#
testing
#
productivity
Comments
1
comment
4 min read
Don't Wire a Coding Model Into Your Workflow Until It Passes Your Own Harness
Dakota Lin
Dakota Lin
Dakota Lin
Follow
Aug 10
Don't Wire a Coding Model Into Your Workflow Until It Passes Your Own Harness
#
ai
#
testing
#
programming
#
tutorial
Comments
Add Comment
5 min read
Hype Cycles Don't Ship My Code: How I Gate New LLMs with a Self-Written Eval Deck
Riley Wu
Riley Wu
Riley Wu
Follow
Aug 10
Hype Cycles Don't Ship My Code: How I Gate New LLMs with a Self-Written Eval Deck
#
ai
#
testing
#
llm
#
productivity
Comments
Add Comment
5 min read
A Two-Model Regression Harness for Evaluating a New Low-Cost Model Release
Avery Wang
Avery Wang
Avery Wang
Follow
Aug 14
A Two-Model Regression Harness for Evaluating a New Low-Cost Model Release
#
ai
#
programming
#
testing
#
productivity
Comments
Add Comment
5 min read
I Stopped Trusting Demo Prompts: A Repeatable Smoke Test for Free Coding Models
Dakota Huang
Dakota Huang
Dakota Huang
Follow
Aug 13
I Stopped Trusting Demo Prompts: A Repeatable Smoke Test for Free Coding Models
#
ai
#
testing
#
productivity
#
tutorial
Comments
Add Comment
5 min read
A New MiniMax Model Dropped: How to Evaluate It on Your Own Code Before Believing the Benchmarks
Avery Wang
Avery Wang
Avery Wang
Follow
Aug 10
A New MiniMax Model Dropped: How to Evaluate It on Your Own Code Before Believing the Benchmarks
#
ai
#
opensource
#
llm
#
programming
Comments
Add Comment
4 min read
A New Cheap Model Dropped. Here's the 2-Hour Canary Test I Run Before Touching It
Jordan Huang
Jordan Huang
Jordan Huang
Follow
Aug 13
A New Cheap Model Dropped. Here's the 2-Hour Canary Test I Run Before Touching It
#
ai
#
llm
#
testing
#
productivity
1
reaction
Comments
1
comment
6 min read
A Reproducible Harness for Evaluating Free AI Coding Models on Your Own Codebase
Dakota Liu
Dakota Liu
Dakota Liu
Follow
Aug 10
A Reproducible Harness for Evaluating Free AI Coding Models on Your Own Codebase
#
ai
#
programming
#
productivity
#
tutorial
Comments
Add Comment
4 min read
Every New Model Gets the Same Six Questions From Me
Dakota Huang
Dakota Huang
Dakota Huang
Follow
Aug 10
Every New Model Gets the Same Six Questions From Me
#
ai
#
testing
#
python
#
productivity
Comments
Add Comment
5 min read
A Practical Scorecard for Deciding If a Free Coding Model Earns a Place in Your Workflow
Taylor Lin
Taylor Lin
Taylor Lin
Follow
Aug 10
A Practical Scorecard for Deciding If a Free Coding Model Earns a Place in Your Workflow
#
ai
#
productivity
#
testing
#
programming
Comments
Add Comment
5 min read
A Personal Scorecard for Evaluating New Coding Models (With the Scripts I Use)
Jordan Huang
Jordan Huang
Jordan Huang
Follow
Aug 10
A Personal Scorecard for Evaluating New Coding Models (With the Scripts I Use)
#
ai
#
programming
#
opensource
#
productivity
Comments
Add Comment
6 min read
1
2
3
4
5
Next ›
Last »
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account