DEV Community

Patrick Hughes profile picture

Patrick Hughes

404 bio not found

Joined Joined on  github website
Ollama num_batch: 256 Was My RTX 5090 Sweet Spot

Ollama num_batch: 256 Was My RTX 5090 Sweet Spot

Comments
4 min read

Want to connect with Patrick Hughes?

Create an account to connect with Patrick Hughes. You can also sign in below to proceed if you already have an account.

Already have an account? Sign in
A 7 GB 27B Model Lost to My 17 GB Default

A 7 GB 27B Model Lost to My 17 GB Default

1
Comments
5 min read
My Agents Have to Prove What They Did

My Agents Have to Prove What They Did

Comments 1
4 min read
Your Agent's Audit Trail Cannot Be Retrofitted

Your Agent's Audit Trail Cannot Be Retrofitted

Comments
5 min read
Ollama Raised $65M. What Builders Get

Ollama Raised $65M. What Builders Get

1
Comments
5 min read
Why production AI is moving to open weights

Why production AI is moving to open weights

Comments
4 min read
Devlog 2026-07-13: local drive git corruption stalls the q

Devlog 2026-07-13: local drive git corruption stalls the q

Comments
3 min read
My 8B Model Failed a 400-Word Task

My 8B Model Failed a 400-Word Task

Comments
5 min read
Q4km vs Q5km: Q4_K_M vs Q5_K_M

Q4km vs Q5km: Q4_K_M vs Q5_K_M

Comments
3 min read
I Rebuilt 5,383 Embeddings After a Dimension Change

I Rebuilt 5,383 Embeddings After a Dimension Change

Comments
5 min read
I Raised the QA Bar on Blog Infographics

I Raised the QA Bar on Blog Infographics

Comments
4 min read
Pin Your Context Window or Pay the Reload Tax

Pin Your Context Window or Pay the Reload Tax

Comments
5 min read
Preview Retrieval Before Your Local LLM Runs

Preview Retrieval Before Your Local LLM Runs

Comments
5 min read
Pin Your Local LLM Context Size Before You Build a Router

Pin Your Local LLM Context Size Before You Build a Router

Comments
4 min read
When Claude hits a weekly limit, your agent fleet still needs a third CLI

When Claude hits a weekly limit, your agent fleet still needs a third CLI

Comments
4 min read
I built a self-improving code model on one RTX 5090. Here is what actually worked.

I built a self-improving code model on one RTX 5090. Here is what actually worked.

Comments
5 min read
Will That Local Model Fit? Do the VRAM Math First

Will That Local Model Fit? Do the VRAM Math First

Comments
5 min read
Local LLMs Need a Timeout Before They Need a Bigger Model

Local LLMs Need a Timeout Before They Need a Bigger Model

Comments
5 min read
Your local LLM is not a worse Claude. It is a different tool.

Your local LLM is not a worse Claude. It is a different tool.

Comments
5 min read
How I Gate a Local Coding Model Before I Trust It

How I Gate a Local Coding Model Before I Trust It

Comments
5 min read
A verifier loop beats a faster local model

A verifier loop beats a faster local model

Comments
4 min read
How I Make Local Model Runs Fail Safely On A 5090

How I Make Local Model Runs Fail Safely On A 5090

Comments
5 min read
How to Make a Local QLoRA Starter Fail Safely

How to Make a Local QLoRA Starter Fail Safely

Comments
5 min read
Use Owner Gates and AgentGuard to Keep AI Agents Moving

Use Owner Gates and AgentGuard to Keep AI Agents Moving

Comments
5 min read
How to Run Local LLM Verifier Loops on Owned Hardware

How to Run Local LLM Verifier Loops on Owned Hardware

Comments
5 min read
AI Agent Memory: What Actually Works in 2026

AI Agent Memory: What Actually Works in 2026

Comments
4 min read
What Anthropic's MITRE ATT&CK report means for solo AI builders

What Anthropic's MITRE ATT&CK report means for solo AI builders

Comments
4 min read
AI Agent Memory in 2026: How It Works and When to Use It

AI Agent Memory in 2026: How It Works and When to Use It

Comments 1
2 min read
Your AI Agent Says "Done." Make It Prove It.

Your AI Agent Says "Done." Make It Prove It.

2
Comments 2
4 min read
Give Your AI Agents an Append-Only Event Log

Give Your AI Agents an Append-Only Event Log

1
Comments
4 min read
A self-healing system can't heal an empty queue

A self-healing system can't heal an empty queue

Comments
4 min read
Missing AI agent cost data is not zero

Missing AI agent cost data is not zero

Comments 1
4 min read
Anthropic Writes 80% of Its Code with Claude

Anthropic Writes 80% of Its Code with Claude

Comments
2 min read
What Salesforce's 20,000 AI Agent Deployments Teach a Solo Builder

What Salesforce's 20,000 AI Agent Deployments Teach a Solo Builder

Comments 2
4 min read
57-71% of AI agents leak data between users. Here's what to do.

57-71% of AI agents leak data between users. Here's what to do.

Comments
3 min read
VRAM Calculator: Estimate Local LLM Requirements

VRAM Calculator: Estimate Local LLM Requirements

Comments
1 min read
Anthropic's IPO and the 40% Cost-Savings Gap: Why Your Spend Cap Matters More Now

Anthropic's IPO and the 40% Cost-Savings Gap: Why Your Spend Cap Matters More Now

Comments
4 min read
When JPMorgan's AI bill goes up, who controls it?

When JPMorgan's AI bill goes up, who controls it?

Comments
4 min read
57-71% of AI Agents Leak Data Between Users. Here's the Fix.

57-71% of AI Agents Leak Data Between Users. Here's the Fix.

Comments
4 min read
AI Coding Assistant Pricing in 2026: Copilot vs Cursor vs Claude Code

AI Coding Assistant Pricing in 2026: Copilot vs Cursor vs Claude Code

Comments
3 min read
How to Close the AI Agent Cost Gap at the Call Site

How to Close the AI Agent Cost Gap at the Call Site

Comments
4 min read
Agentic coding moved my bottleneck to code review

Agentic coding moved my bottleneck to code review

Comments
3 min read
How to Pick a GGUF Quant Level for Your VRAM Budget

How to Pick a GGUF Quant Level for Your VRAM Budget

Comments
4 min read
Stop Telling People You Have 11 AI Agents

Stop Telling People You Have 11 AI Agents

Comments
4 min read
GGUF Quantization and VRAM: How to Pick Q4, Q5, or Q8 for Your GPU (2026)

GGUF Quantization and VRAM: How to Pick Q4, Q5, or Q8 for Your GPU (2026)

Comments
4 min read
Your AI agent doesn't need memory. It needs a file.

Your AI agent doesn't need memory. It needs a file.

1
Comments
4 min read
How to Tune llama.cpp --n-gpu-layers: A Practical VRAM Guide (2026)

How to Tune llama.cpp --n-gpu-layers: A Practical VRAM Guide (2026)

Comments
4 min read
Which GGUF Quant Should You Actually Pick? Q4 vs Q5 vs Q6 vs Q8 (2026)

Which GGUF Quant Should You Actually Pick? Q4 vs Q5 vs Q6 vs Q8 (2026)

Comments
4 min read
How to Tune --n-gpu-layers for Your VRAM Budget

How to Tune --n-gpu-layers for Your VRAM Budget

Comments
4 min read
llama.cpp Multi-GPU: Splitting a Model Across Cards with --tensor-split

llama.cpp Multi-GPU: Splitting a Model Across Cards with --tensor-split

Comments
5 min read
What Uber's $1,500/Developer AI Cap Tells You About Your Own Bill

What Uber's $1,500/Developer AI Cap Tells You About Your Own Bill

Comments
4 min read
Your AI Agent's Retry Loop Is a Cost Bug Waiting to Happen

Your AI Agent's Retry Loop Is a Cost Bug Waiting to Happen

Comments
3 min read
When JPMorgan Turns On AI Bank-Wide, Who Controls the Bill?

When JPMorgan Turns On AI Bank-Wide, Who Controls the Bill?

Comments
4 min read
What Anthropic's MITRE ATT&CK Report Means for Teams Running AI Agents

What Anthropic's MITRE ATT&CK Report Means for Teams Running AI Agents

Comments
4 min read
What GitHub Copilot Users Wish They Had a Week Ago

What GitHub Copilot Users Wish They Had a Week Ago

Comments
3 min read
When Not to Use an AI Agent

When Not to Use an AI Agent

Comments
3 min read
llama.cpp ngl: when -ngl 99 still runs on your CPU

llama.cpp ngl: when -ngl 99 still runs on your CPU

1
Comments
5 min read
I made my blog API reject its own writer

I made my blog API reject its own writer

Comments
4 min read
When Your Blog Repair Loop Fails 23 Times, Stop Repairing

When Your Blog Repair Loop Fails 23 Times, Stop Repairing

Comments
3 min read
Your Cron Jobs Lie - Why I Built an Outcome Checker

Your Cron Jobs Lie - Why I Built an Outcome Checker

1
Comments
4 min read
loading...