DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
AI Metrics Baseline: Prove Your Feature Works Before Scaling It

AI Metrics Baseline: Prove Your Feature Works Before Scaling It

1
Comments
9 min read
Cross-Harness Tool Parity: Write One Custom MCP Tool, Deploy It Everywhere

Cross-Harness Tool Parity: Write One Custom MCP Tool, Deploy It Everywhere

Comments
5 min read
AirLLM Runs a 70B Model on a 4GB GPU. It's True, and That's Not the Interesting Part

AirLLM Runs a 70B Model on a 4GB GPU. It's True, and That's Not the Interesting Part

7
Comments 3
7 min read
If AI writes code, what is our job now?

If AI writes code, what is our job now?

Comments
2 min read
Spring AI Prompt Caching and Chat Memory: Where the Tokens Go — LLM Cost Control 2/4

Spring AI Prompt Caching and Chat Memory: Where the Tokens Go — LLM Cost Control 2/4

Comments 2
10 min read
The bug that took me four hours to find had nothing to do with the model

The bug that took me four hours to find had nothing to do with the model

1
Comments 1
2 min read
Your AI agent can refuse to leak a secret — and leak it anyway, in its "thinking"

Your AI agent can refuse to leak a secret — and leak it anyway, in its "thinking"

1
Comments
6 min read
# Securing API Tokens: Protecting Your AI Applications from Credential Leakage

# Securing API Tokens: Protecting Your AI Applications from Credential Leakage

Comments
4 min read
Beyond Logs: Building a Real-Time AI Observability Dashboard That Surfaces Database Rows, Not Just Latency Percentiles

Beyond Logs: Building a Real-Time AI Observability Dashboard That Surfaces Database Rows, Not Just Latency Percentiles

Comments
5 min read
Browser Harness hands the LLM a websocket to Chrome and lets it write the rest

Browser Harness hands the LLM a websocket to Chrome and lets it write the rest

1
Comments
3 min read
Stop AI Agent Drift Across Sessions With Versioned, Grep-able Rules

Stop AI Agent Drift Across Sessions With Versioned, Grep-able Rules

2
Comments 2
5 min read
Google's V2A is the other half of generative video

Google's V2A is the other half of generative video

Comments
3 min read
Deploying Large Language Models (LLMs) with Ollama

Deploying Large Language Models (LLMs) with Ollama

1
Comments
3 min read
Never trust an LLM's output directly. Here's the validation layer I put on every agent.

Never trust an LLM's output directly. Here's the validation layer I put on every agent.

Comments
5 min read
Context rot: why your AI agent gets dumber the longer it runs

Context rot: why your AI agent gets dumber the longer it runs

Comments
6 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.