DEV Community

Aamer Mihaysi profile picture

Aamer Mihaysi

AI Engineer building agentic systems, RAG pipelines & multi-agent swarms. LLM tuning, production MLOps, open-source first. Making AI practical and deployable.

Location Qatar, Doha Joined Joined on  Personal website https://www.mehaisi.com github website
How much of your agent's sandbox is actually read-only?

How much of your agent's sandbox is actually read-only?

Comments
3 min read
Four new models, one interface change: from chat completion to decision

Four new models, one interface change: from chat completion to decision

Comments
5 min read
Blast radius is the unit of trust

Blast radius is the unit of trust

Comments 1
5 min read
My World-Model Drift Metric Scored a Stale Checkpoint as Stable

My World-Model Drift Metric Scored a Stale Checkpoint as Stable

Comments
4 min read
Prune tool output by rule, leave the reasoning chain alone

Prune tool output by rule, leave the reasoning chain alone

Comments
4 min read
Prompt search is a hill-climber, and accuracy is the wrong hill

Prompt search is a hill-climber, and accuracy is the wrong hill

Comments
4 min read
Uniform INT8 on Recurrent States Is a Default, Not a Decision

Uniform INT8 on Recurrent States Is a Default, Not a Decision

Comments
4 min read
If every layer prefix is a valid model, why do we still pick a size at deploy time?

If every layer prefix is a valid model, why do we still pick a size at deploy time?

Comments
5 min read
The agent didn't fail. You never told it what "done" means.

The agent didn't fail. You never told it what "done" means.

Comments
2 min read
Containers aren't a sandbox: the Fedora agent incident

Containers aren't a sandbox: the Fedora agent incident

Comments
4 min read
The Windows 11 agent runs in the background, and that's the part nobody's talking about

The Windows 11 agent runs in the background, and that's the part nobody's talking about

Comments
2 min read
The benchmark isn't the problem. Your architecture is.

The benchmark isn't the problem. Your architecture is.

Comments
2 min read
The agent has no skin in the game. That's the bug.

The agent has no skin in the game. That's the bug.

1
Comments 2
2 min read
Your agent deleted prod because you never let it fail in staging

Your agent deleted prod because you never let it fail in staging

1
Comments 1
3 min read
The agent runtime is the new IDE. The audit trail is the new GitHub.

The agent runtime is the new IDE. The audit trail is the new GitHub.

Comments
4 min read
The model that argued with my prompt

The model that argued with my prompt

Comments
4 min read
The sandbox can't see what's already in the context

The sandbox can't see what's already in the context

Comments
3 min read
Your agent isn't lying. It's optimizing.

Your agent isn't lying. It's optimizing.

Comments
4 min read
The person on the other end of your agent's output

The person on the other end of your agent's output

Comments
2 min read
You gave your agent maintainer access. That's the bug.

You gave your agent maintainer access. That's the bug.

Comments
2 min read
The Research Bottleneck Isn't Ideas — It's the Missing 40% of the Paper

The Research Bottleneck Isn't Ideas — It's the Missing 40% of the Paper

Comments
3 min read
The 30–50% violation rate is a baseline, not a verdict

The 30–50% violation rate is a baseline, not a verdict

Comments
2 min read
The agent that shamed a maintainer wasn't rude. It was optimized.

The agent that shamed a maintainer wasn't rude. It was optimized.

1
Comments
4 min read
The real scandal isn't the benchmark gaming. It's the pretending.

The real scandal isn't the benchmark gaming. It's the pretending.

Comments
2 min read
Open-source doesn't mean safe: the guardrail check for coding agents

Open-source doesn't mean safe: the guardrail check for coding agents

Comments
4 min read
Sandboxes contain the blast radius, not the consequences

Sandboxes contain the blast radius, not the consequences

Comments
3 min read
Your LLM Judge Can't See What Feedback Did — and It's Costing You

Your LLM Judge Can't See What Feedback Did — and It's Costing You

Comments
4 min read
Your agent can drop the database. Who approves that?

Your agent can drop the database. Who approves that?

Comments
4 min read
The DN42 agent didn't need a budget. It needed a sense of cost.

The DN42 agent didn't need a budget. It needed a sense of cost.

Comments
2 min read
The skill bottleneck is a myth — your agent needs a memory layer

The skill bottleneck is a myth — your agent needs a memory layer

1
Comments
4 min read
SWE-Prime: the pass label is a terrible filter for agent training data

SWE-Prime: the pass label is a terrible filter for agent training data

Comments
4 min read
The reward function is a policy document

The reward function is a policy document

Comments
3 min read
Benchmark scores are marketing now

Benchmark scores are marketing now

1
Comments
4 min read
The proof kernel is the only autonomy you get

The proof kernel is the only autonomy you get

Comments
4 min read
Your agent's memory doesn't live in the sandbox

Your agent's memory doesn't live in the sandbox

Comments
3 min read
The matplotlib PR isn't an agent problem. It's a moderation problem.

The matplotlib PR isn't an agent problem. It's a moderation problem.

Comments
2 min read
Your agent passes SWE-bench. It still can't do science.

Your agent passes SWE-bench. It still can't do science.

Comments
4 min read
The agent wrote a hit piece because you asked it to

The agent wrote a hit piece because you asked it to

1
Comments
2 min read
The default is unlimited

The default is unlimited

1
Comments
2 min read
Your OS agent should not have your keys

Your OS agent should not have your keys

Comments
4 min read
OpenCode vs Cursor vs Copilot: the model was never the bottleneck

OpenCode vs Cursor vs Copilot: the model was never the bottleneck

1
Comments
4 min read
The reward signal is the bottleneck, not the model

The reward signal is the bottleneck, not the model

Comments
3 min read
Selection pressure turns benchmarks into fingerprints

Selection pressure turns benchmarks into fingerprints

Comments
4 min read
Who teaches an agent to take the L?

Who teaches an agent to take the L?

Comments 1
4 min read
Sandboxes don't stop the bill

Sandboxes don't stop the bill

Comments
4 min read
The Career Isn't Eroding. It's Relocating.

The Career Isn't Eroding. It's Relocating.

1
Comments
2 min read
Bonsai-27B on a Single 3090: What Works and What Doesn't

Bonsai-27B on a Single 3090: What Works and What Doesn't

1
Comments
3 min read
By Mid-2027 You'll Train LLMs From Scratch, Not Fine-Tune

By Mid-2027 You'll Train LLMs From Scratch, Not Fine-Tune

1
Comments
2 min read
You Don't Need a Better Agent. You Need a Better Debug Loop.

You Don't Need a Better Agent. You Need a Better Debug Loop.

1
Comments
2 min read
Fifty Poisoned Samples Is All It Takes

Fifty Poisoned Samples Is All It Takes

Comments
3 min read
Small Models That Think Harder Beat Big Models That Sound Confident

Small Models That Think Harder Beat Big Models That Sound Confident

Comments
2 min read
Llamafile vs vLLM: Two Ways to Serve a Local Model, and When Each Makes Sense

Llamafile vs vLLM: Two Ways to Serve a Local Model, and When Each Makes Sense

Comments 1
3 min read
Most Evals Measure the Wrong Thing

Most Evals Measure the Wrong Thing

Comments
2 min read
Agent sandboxing will be the default by mid-2027, and most teams aren't ready

Agent sandboxing will be the default by mid-2027, and most teams aren't ready

5
Comments 1
3 min read
GLM-5.2: The 1M-Context Open Model That Actually Ships

GLM-5.2: The 1M-Context Open Model That Actually Ships

Comments
3 min read
Leanstral 1.5: Mistral's Open-Source Bet on Formal Verification Agents

Leanstral 1.5: Mistral's Open-Source Bet on Formal Verification Agents

Comments
3 min read
The Reality of 1M Context: Testing Qwythos-9B-Claude-Mythos-5

The Reality of 1M Context: Testing Qwythos-9B-Claude-Mythos-5

Comments
2 min read
The Persistence Trap: Why Autonomous Agents are a New Security Nightmare

The Persistence Trap: Why Autonomous Agents are a New Security Nightmare

1
Comments 1
3 min read
Stop Treating LLM Memory as a Database: The Shift Toward Memory as a Skill

Stop Treating LLM Memory as a Database: The Shift Toward Memory as a Skill

1
Comments
3 min read
Stop Chasing the Hype: Real-World Takeaways from DeepSeek-V4-Pro-DSpark

Stop Chasing the Hype: Real-World Takeaways from DeepSeek-V4-Pro-DSpark

1
Comments
2 min read
loading...