DEV Community

#llm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%

Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%

Comments
5 min read
Your MCP Server's Client Is a Language Model. Write the Contract Like You Mean It

Your MCP Server's Client Is a Language Model. Write the Contract Like You Mean It

1
Comments
4 min read
Four Ways to Build Solon AI Tools: Annotations, FunctionToolDesc, returnDirect, and toolContext

Four Ways to Build Solon AI Tools: Annotations, FunctionToolDesc, returnDirect, and toolContext

Comments
6 min read
Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics

Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics

Comments
5 min read
Running Local LLMs in Java: Introducing jllm – A Minimalist Ollama Alternative

Running Local LLMs in Java: Introducing jllm – A Minimalist Ollama Alternative

Comments
3 min read
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%

Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%

Comments
5 min read
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%

Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%

Comments
5 min read
Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics

Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics

Comments
5 min read
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%

Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%

Comments
5 min read
Your cheap LLM relay might be swapping the model. Here's how to catch it.

Your cheap LLM relay might be swapping the model. Here's how to catch it.

Comments
7 min read
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%

Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%

Comments
5 min read
Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%

Optimizing RAG at Scale: Chunking, Retrieval, and the Bayesian Search That Cut Latency 40%

Comments
5 min read
Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics

Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics

Comments
5 min read
Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics

Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics

Comments
5 min read
Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics

Building Production-Grade LLM Evaluation Pipelines: From Vibes to Metrics

Comments 1
5 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.