DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
The Model Scored 30%. The Harness Scored 100%. Which One Did You Benchmark?

Harnesses drive massive score jumps

The Model Scored 30%. The Harness Scored 100%. Which One Did You Benchmark?

22
Comments 44
10 min read
AI Weekly: GPT-5.6-Cyber, Muse Glimmer, and the Agent Browser

AI Weekly: GPT-5.6-Cyber, Muse Glimmer, and the Agent Browser

1
Comments 1
24 min read
Challenging Tokenization: An LLM Architecture Experiment Without a Tokenizer

Challenging Tokenization: An LLM Architecture Experiment Without a Tokenizer

Comments
3 min read
Probe vs Prose: what the verifier-sharing-your-text-channel really costs

Probe vs Prose: what the verifier-sharing-your-text-channel really costs

4
Comments 2
16 min read
AI This Week (Aug 2026): Qwen3.8 Max, DeepSeek V4-Flash, and Models Shipping Like Patches

AI This Week (Aug 2026): Qwen3.8 Max, DeepSeek V4-Flash, and Models Shipping Like Patches

Comments
2 min read
300+ languages doesn't mean what you think: three tiers of code understanding

300+ languages doesn't mean what you think: three tiers of code understanding

Comments
7 min read
Keeping the LLM out of the verdict

Keeping the LLM out of the verdict

Comments 1
5 min read
How parallel AI agents should talk to each other (and the bug that proved it)

How parallel AI agents should talk to each other (and the bug that proved it)

1
Comments 2
3 min read
I tried to patch a blind spot in my own MCP tool. The patch and a false-positive bug cancel out.

I tried to patch a blind spot in my own MCP tool. The patch and a false-positive bug cancel out.

2
Comments
5 min read
I built an async wrapper for OpenAI/Anthropic SDKs because I didn't want a proxy in my request path

I built an async wrapper for OpenAI/Anthropic SDKs because I didn't want a proxy in my request path

Comments
2 min read
I created a tool to inspect full python codebases

I created a tool to inspect full python codebases

Comments
1 min read
When your benchmark is wrong and your model is right

When your benchmark is wrong and your model is right

1
Comments
7 min read
What changed in Apiarium after developers started using it

What changed in Apiarium after developers started using it

19
Comments 3
3 min read
"How to Tell If an LLM Was Really Trained From Scratch: A Reproducible Fingerprinting Method"

"How to Tell If an LLM Was Really Trained From Scratch: A Reproducible Fingerprinting Method"

Comments
7 min read
🤖 AI Context Engineering (Part 4): AI Agents - From Tool Calling to Multi-Step Workflows

🤖 AI Context Engineering (Part 4): AI Agents - From Tool Calling to Multi-Step Workflows

1
Comments
9 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.