DEV Community

Aamer Mihaysi profile picture

Aamer Mihaysi

AI Engineer building agentic systems, RAG pipelines & multi-agent swarms. LLM tuning, production MLOps, open-source first. Making AI practical and deployable.

Location Qatar, Doha Joined Joined on  Personal website https://www.mehaisi.com github website
The agent runtime is the new IDE. The audit trail is the new GitHub.

The agent runtime is the new IDE. The audit trail is the new GitHub.

Comments
4 min read
The model that argued with my prompt

The model that argued with my prompt

Comments
4 min read
The sandbox can't see what's already in the context

The sandbox can't see what's already in the context

Comments
3 min read
Your agent isn't lying. It's optimizing.

Your agent isn't lying. It's optimizing.

Comments
4 min read
The person on the other end of your agent's output

The person on the other end of your agent's output

Comments
2 min read
You gave your agent maintainer access. That's the bug.

You gave your agent maintainer access. That's the bug.

Comments
2 min read
The Research Bottleneck Isn't Ideas — It's the Missing 40% of the Paper

The Research Bottleneck Isn't Ideas — It's the Missing 40% of the Paper

Comments
3 min read
The 30–50% violation rate is a baseline, not a verdict

The 30–50% violation rate is a baseline, not a verdict

Comments
2 min read
The agent that shamed a maintainer wasn't rude. It was optimized.

The agent that shamed a maintainer wasn't rude. It was optimized.

1
Comments
4 min read
The real scandal isn't the benchmark gaming. It's the pretending.

The real scandal isn't the benchmark gaming. It's the pretending.

Comments
2 min read
Open-source doesn't mean safe: the guardrail check for coding agents

Open-source doesn't mean safe: the guardrail check for coding agents

Comments
4 min read
Sandboxes contain the blast radius, not the consequences

Sandboxes contain the blast radius, not the consequences

Comments
3 min read
Your LLM Judge Can't See What Feedback Did — and It's Costing You

Your LLM Judge Can't See What Feedback Did — and It's Costing You

Comments
4 min read
Your agent can drop the database. Who approves that?

Your agent can drop the database. Who approves that?

Comments
4 min read
The DN42 agent didn't need a budget. It needed a sense of cost.

The DN42 agent didn't need a budget. It needed a sense of cost.

Comments
2 min read
The skill bottleneck is a myth — your agent needs a memory layer

The skill bottleneck is a myth — your agent needs a memory layer

1
Comments
4 min read
SWE-Prime: the pass label is a terrible filter for agent training data

SWE-Prime: the pass label is a terrible filter for agent training data

Comments
4 min read
The reward function is a policy document

The reward function is a policy document

Comments
3 min read
Benchmark scores are marketing now

Benchmark scores are marketing now

1
Comments
4 min read
The proof kernel is the only autonomy you get

The proof kernel is the only autonomy you get

Comments
4 min read
Your agent's memory doesn't live in the sandbox

Your agent's memory doesn't live in the sandbox

Comments
3 min read
The matplotlib PR isn't an agent problem. It's a moderation problem.

The matplotlib PR isn't an agent problem. It's a moderation problem.

Comments
2 min read
Your agent passes SWE-bench. It still can't do science.

Your agent passes SWE-bench. It still can't do science.

Comments
4 min read
The agent wrote a hit piece because you asked it to

The agent wrote a hit piece because you asked it to

1
Comments
2 min read
The default is unlimited

The default is unlimited

1
Comments
2 min read
Your OS agent should not have your keys

Your OS agent should not have your keys

Comments
4 min read
OpenCode vs Cursor vs Copilot: the model was never the bottleneck

OpenCode vs Cursor vs Copilot: the model was never the bottleneck

1
Comments
4 min read
The reward signal is the bottleneck, not the model

The reward signal is the bottleneck, not the model

Comments
3 min read
Selection pressure turns benchmarks into fingerprints

Selection pressure turns benchmarks into fingerprints

Comments
4 min read
Who teaches an agent to take the L?

Who teaches an agent to take the L?

Comments 1
4 min read
Sandboxes don't stop the bill

Sandboxes don't stop the bill

Comments
4 min read
The Career Isn't Eroding. It's Relocating.

The Career Isn't Eroding. It's Relocating.

1
Comments
2 min read
Bonsai-27B on a Single 3090: What Works and What Doesn't

Bonsai-27B on a Single 3090: What Works and What Doesn't

1
Comments
3 min read
By Mid-2027 You'll Train LLMs From Scratch, Not Fine-Tune

By Mid-2027 You'll Train LLMs From Scratch, Not Fine-Tune

1
Comments
2 min read
You Don't Need a Better Agent. You Need a Better Debug Loop.

You Don't Need a Better Agent. You Need a Better Debug Loop.

1
Comments
2 min read
Fifty Poisoned Samples Is All It Takes

Fifty Poisoned Samples Is All It Takes

Comments
3 min read
Small Models That Think Harder Beat Big Models That Sound Confident

Small Models That Think Harder Beat Big Models That Sound Confident

Comments
2 min read
Llamafile vs vLLM: Two Ways to Serve a Local Model, and When Each Makes Sense

Llamafile vs vLLM: Two Ways to Serve a Local Model, and When Each Makes Sense

Comments 1
3 min read
Most Evals Measure the Wrong Thing

Most Evals Measure the Wrong Thing

Comments
2 min read
Agent sandboxing will be the default by mid-2027, and most teams aren't ready

Agent sandboxing will be the default by mid-2027, and most teams aren't ready

5
Comments 1
3 min read
GLM-5.2: The 1M-Context Open Model That Actually Ships

GLM-5.2: The 1M-Context Open Model That Actually Ships

Comments
3 min read
Leanstral 1.5: Mistral's Open-Source Bet on Formal Verification Agents

Leanstral 1.5: Mistral's Open-Source Bet on Formal Verification Agents

Comments
3 min read
The Reality of 1M Context: Testing Qwythos-9B-Claude-Mythos-5

The Reality of 1M Context: Testing Qwythos-9B-Claude-Mythos-5

Comments
2 min read
The Persistence Trap: Why Autonomous Agents are a New Security Nightmare

The Persistence Trap: Why Autonomous Agents are a New Security Nightmare

1
Comments 1
3 min read
Stop Treating LLM Memory as a Database: The Shift Toward Memory as a Skill

Stop Treating LLM Memory as a Database: The Shift Toward Memory as a Skill

1
Comments
3 min read
Stop Chasing the Hype: Real-World Takeaways from DeepSeek-V4-Pro-DSpark

Stop Chasing the Hype: Real-World Takeaways from DeepSeek-V4-Pro-DSpark

1
Comments
2 min read
Testing Qwen-AgentWorld-35B-A3B: A New Benchmark for Agentic Reasoning?

Testing Qwen-AgentWorld-35B-A3B: A New Benchmark for Agentic Reasoning?

Comments
2 min read
Why 1M Context Windows Actually Matter: Testing Qwythos-9B-Claude-Mythos

Why 1M Context Windows Actually Matter: Testing Qwythos-9B-Claude-Mythos

Comments
3 min read
Gemma 4 12B Coder: The New Sweet Spot for Local Agentic Workflows

Gemma 4 12B Coder: The New Sweet Spot for Local Agentic Workflows

Comments
2 min read
Local Agents: Why Gemma 4 12B Agentic is the Sweet Spot for Production

Local Agents: Why Gemma 4 12B Agentic is the Sweet Spot for Production

Comments
2 min read
Beyond the Hype: Testing Gemma-4-12B Agentic GGUFs in the Wild

Beyond the Hype: Testing Gemma-4-12B Agentic GGUFs in the Wild

Comments
2 min read
Async Batching Is the Real Latency Win Nobody's Talking About

Async Batching Is the Real Latency Win Nobody's Talking About

1
Comments 1
3 min read
DeepSeek-V4: Finally, a Context Window Built for Agents

DeepSeek-V4: Finally, a Context Window Built for Agents

2
Comments 2
2 min read
EMO: Mixture-of-Experts That Actually Behaves Like One

EMO: Mixture-of-Experts That Actually Behaves Like One

1
Comments 1
3 min read
TPUs for the Agentic Era: Hardware Finally Catching Up to the Workload

TPUs for the Agentic Era: Hardware Finally Catching Up to the Workload

Comments
2 min read
MoE Architectures Keep Solving the Wrong Problem

MoE Architectures Keep Solving the Wrong Problem

Comments
3 min read
MachinaCheck: Manufacturing Agents That Actually Ship

MachinaCheck: Manufacturing Agents That Actually Ship

Comments
2 min read
vLLM's V1 Release Fixes the Silent Killer in RL Training

vLLM's V1 Release Fixes the Silent Killer in RL Training

Comments
2 min read
DeepSeek-V4: What a Million-Token Context Actually Changes

DeepSeek-V4: What a Million-Token Context Actually Changes

Comments 1
3 min read
The 8B Model That Punches at 32B Weight

The 8B Model That Punches at 32B Weight

Comments
2 min read
loading...