DEV Community

Patrick Hughes profile picture

Patrick Hughes

404 bio not found

Joined Joined on  github website
A 32 GB GPU Still Needs Host RAM Headroom

A 32 GB GPU Still Needs Host RAM Headroom

Comments
5 min read

Want to connect with Patrick Hughes?

Create an account to connect with Patrick Hughes. You can also sign in below to proceed if you already have an account.

Already have an account? Sign in
Your Benchmark Row Never Saved the Driver Version

Your Benchmark Row Never Saved the Driver Version

Comments
5 min read
Your Local LLM CSV Needs a Schema Version

Your Local LLM CSV Needs a Schema Version

Comments
5 min read
Preflight Local AI Before You Benchmark a Model

Preflight Local AI Before You Benchmark a Model

Comments
5 min read
49W Average Hid a 338W Burst on Gemma 26B

49W Average Hid a 338W Burst on Gemma 26B

Comments
5 min read
How to Calculate Local LLM Energy per Token

How to Calculate Local LLM Energy per Token

Comments
5 min read
Why a Failed Local LLM Benchmark Row Still Matters

Why a Failed Local LLM Benchmark Row Still Matters

Comments
4 min read
My local LLM eval hid four token caps

My local LLM eval hid four token caps

Comments
4 min read
Mojo Is Open Source. What Local AI Builders Need to Know

Mojo Is Open Source. What Local AI Builders Need to Know

Comments
5 min read
A GPU Driver Is Not a Local LLM Benchmark

A GPU Driver Is Not a Local LLM Benchmark

Comments
4 min read
The 26B Model Hit the Cap. The 8B Finished.

The 26B Model Hit the Cap. The 8B Finished.

Comments
4 min read
My Local LLM Writer Failed Its Own Word-Count Gate

My Local LLM Writer Failed Its Own Word-Count Gate

Comments
4 min read
A 27B Model Fit on an 8 GB GPU. It Was Slow.

A 27B Model Fit on an 8 GB GPU. It Was Slow.

Comments
5 min read
Unload Local LLMs After Every Test

Unload Local LLMs After Every Test

Comments
4 min read
How I Test a 30B Local Model Before I Load It

How I Test a 30B Local Model Before I Load It

Comments
5 min read
How I Benchmark Local LLMs Before I Trust Them

How I Benchmark Local LLMs Before I Trust Them

Comments
4 min read
When a 4B Local LLM Beats 26B on One Task

When a 4B Local LLM Beats 26B on One Task

Comments
5 min read
Does Ollama Include That New llama.cpp Feature?

Does Ollama Include That New llama.cpp Feature?

Comments
4 min read
My local models refused zero of 50 security tasks

My local models refused zero of 50 security tasks

Comments
5 min read
Prime Agent hit 95.5% on ARC-AGI-3. I did not install it.

Prime Agent hit 95.5% on ARC-AGI-3. I did not install it.

Comments
6 min read
My Local LLM Got Faster After It Passed the Tests

My Local LLM Got Faster After It Passed the Tests

Comments
4 min read
Chunk Size Is a Reliability Setting

Chunk Size Is a Reliability Setting

Comments
5 min read
The faster local model run took 83x longer

The faster local model run took 83x longer

Comments
5 min read
VRAM Fit Is Not Runtime Support

VRAM Fit Is Not Runtime Support

Comments
4 min read
Why Local LLM Benchmarks Need Power Data

Why Local LLM Benchmarks Need Power Data

Comments
4 min read
My 5090 benchmark was missing the field I needed most

My 5090 benchmark was missing the field I needed most

Comments
4 min read
Search Old Results Before Publishing an LLM Test

Search Old Results Before Publishing an LLM Test

Comments
5 min read
Build Local LLM Eval Data From Real Failures

Build Local LLM Eval Data From Real Failures

Comments
5 min read
How I Keep LLM Results Valid After a Driver Update

How I Keep LLM Results Valid After a Driver Update

Comments
4 min read
How I Budget VRAM for Shared Local AI Workloads

How I Budget VRAM for Shared Local AI Workloads

Comments
5 min read
Why I Benchmark Local LLM Input and Output Separately

Why I Benchmark Local LLM Input and Output Separately

Comments
5 min read
Why I Test 3 Workloads Before Sizing a Local LLM

Why I Test 3 Workloads Before Sizing a Local LLM

Comments
5 min read
Incident response needs a local model you already trust

Incident response needs a local model you already trust

Comments
6 min read
Local open-model agents just became a product category

Local open-model agents just became a product category

Comments
6 min read
Your local LLM benchmark is probably lying to you

Your local LLM benchmark is probably lying to you

Comments
5 min read
Test Retrieval Before Your Local LLM Writes

Test Retrieval Before Your Local LLM Writes

Comments
4 min read
Why I Did Not Promote My Smaller Local Model

Why I Did Not Promote My Smaller Local Model

Comments
4 min read
Build an AI Research Workbench on Your Own GPU

Build an AI Research Workbench on Your Own GPU

Comments
5 min read
Ollama num_batch: 256 Was My RTX 5090 Sweet Spot

Ollama num_batch: 256 Was My RTX 5090 Sweet Spot

Comments
4 min read
A 7 GB 27B Model Lost to My 17 GB Default

A 7 GB 27B Model Lost to My 17 GB Default

1
Comments
5 min read
My Agents Have to Prove What They Did

My Agents Have to Prove What They Did

Comments 1
4 min read
Your Agent's Audit Trail Cannot Be Retrofitted

Your Agent's Audit Trail Cannot Be Retrofitted

Comments
5 min read
Ollama Raised $65M. What Builders Get

Ollama Raised $65M. What Builders Get

1
Comments
5 min read
Why production AI is moving to open weights

Why production AI is moving to open weights

Comments
4 min read
Devlog 2026-07-13: local drive git corruption stalls the q

Devlog 2026-07-13: local drive git corruption stalls the q

Comments
3 min read
My 8B Model Failed a 400-Word Task

My 8B Model Failed a 400-Word Task

Comments
5 min read
Q4km vs Q5km: Q4_K_M vs Q5_K_M

Q4km vs Q5km: Q4_K_M vs Q5_K_M

Comments
3 min read
I Rebuilt 5,383 Embeddings After a Dimension Change

I Rebuilt 5,383 Embeddings After a Dimension Change

Comments
5 min read
I Raised the QA Bar on Blog Infographics

I Raised the QA Bar on Blog Infographics

Comments
4 min read
Pin Your Context Window or Pay the Reload Tax

Pin Your Context Window or Pay the Reload Tax

Comments
5 min read
Preview Retrieval Before Your Local LLM Runs

Preview Retrieval Before Your Local LLM Runs

Comments
5 min read
Pin Your Local LLM Context Size Before You Build a Router

Pin Your Local LLM Context Size Before You Build a Router

Comments
4 min read
When Claude hits a weekly limit, your agent fleet still needs a third CLI

When Claude hits a weekly limit, your agent fleet still needs a third CLI

Comments
4 min read
I built a self-improving code model on one RTX 5090. Here is what actually worked.

I built a self-improving code model on one RTX 5090. Here is what actually worked.

Comments
5 min read
Will That Local Model Fit? Do the VRAM Math First

Will That Local Model Fit? Do the VRAM Math First

Comments
5 min read
Local LLMs Need a Timeout Before They Need a Bigger Model

Local LLMs Need a Timeout Before They Need a Bigger Model

Comments
5 min read
Your local LLM is not a worse Claude. It is a different tool.

Your local LLM is not a worse Claude. It is a different tool.

Comments
5 min read
How I Gate a Local Coding Model Before I Trust It

How I Gate a Local Coding Model Before I Trust It

Comments
5 min read
A verifier loop beats a faster local model

A verifier loop beats a faster local model

Comments
4 min read
How I Make Local Model Runs Fail Safely On A 5090

How I Make Local Model Runs Fail Safely On A 5090

Comments
5 min read
loading...