DEV Community

Artificial Intelligence

Artificial intelligence leverages computers and machines to mimic the problem-solving and decision-making capabilities found in humans and in nature.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
I gave an AI agent a production rollback button — then spent the hackathon trying to trick it into pressing it

Unannotated tools silently bypass approval gates

I gave an AI agent a production rollback button — then spent the hackathon trying to trick it into pressing it

28
Comments 17
17 min read
I Built an AI Name Generator Because Naming Is Harder Than It Looks

I Built an AI Name Generator Because Naming Is Harder Than It Looks

Comments
2 min read
A 7.5B model beat a 24B on my coding benchmark.

A 7.5B model beat a 24B on my coding benchmark.

Comments
21 min read
Google Gemini Expands Connected Apps With Dropbox, Zillow and Viator

Google Gemini Expands Connected Apps With Dropbox, Zillow and Viator

Comments
5 min read
When AI Agents Start Moving Money: Who Controls the Wallet?

When AI Agents Start Moving Money: Who Controls the Wallet?

6
Comments 2
3 min read
One API Key for OpenAI, Claude, and Gemini LLM Classification Routing

One API Key for OpenAI, Claude, and Gemini LLM Classification Routing

Comments
5 min read
I built CleanSlate, an open-source coding agent for the IDE, CLI, and SDK

I built CleanSlate, an open-source coding agent for the IDE, CLI, and SDK

Comments
1 min read
I Ran 3 Open-Weight LLMs Head-to-Head on a 24GB Mac — One Was 3x Faster

I Ran 3 Open-Weight LLMs Head-to-Head on a 24GB Mac — One Was 3x Faster

2
Comments 6
5 min read
Spec-driven development with AI agents: constitutions, checkpoints, and handoffs

Spec-driven development with AI agents: constitutions, checkpoints, and handoffs

Comments
5 min read
Benchmarking GPT-4o, Claude 3.5 Sonnet, and Llama 3 for Automated Code Auditing & Vulnerability Detection

Benchmarking GPT-4o, Claude 3.5 Sonnet, and Llama 3 for Automated Code Auditing & Vulnerability Detection

Comments
2 min read
Nobody has shipped an LLM that grades its own draft without a human in the loop

Nobody has shipped an LLM that grades its own draft without a human in the loop

2
Comments
2 min read
My Repo's Riskiest Regex Has Regressed Three Times on Record. It Never Got the Test Pattern I Gave Everything Else.

My Repo's Riskiest Regex Has Regressed Three Times on Record. It Never Got the Test Pattern I Gave Everything Else.

Comments 1
5 min read
I Thought This Was a Classification Problem. It Wasn't.

I Thought This Was a Classification Problem. It Wasn't.

12
Comments
6 min read
False Completion Is the Real Failure Mode of Coding Agents

False Completion Is the Real Failure Mode of Coding Agents

1
Comments 11
7 min read
Open Knowledge Format vs RAG: Why Your Agent Should Read a Wiki

Open Knowledge Format vs RAG: Why Your Agent Should Read a Wiki

3
Comments 2
9 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.