DEV Community

Testing AI Systems Series' Articles

Back to Sara Bezjak's Series
Five ways to test an LLM's answer and what each one misses

Five ways to test an LLM's answer and what each one misses

2
Comments 1
6 min read
Search bug or model bug - testing a RAG system to tell them apart

Search bug or model bug - testing a RAG system to tell them apart

2
Comments 2
8 min read
The hard part of attacking an AI isn't breaking it. It's telling real harm from fake.

The hard part of attacking an AI isn't breaking it. It's telling real harm from fake.

1
Comments
7 min read
With an AI agent, the answer is the last place the bug shows up

With an AI agent, the answer is the last place the bug shows up

8
Comments 4
6 min read
A benchmark is only as good as the model you use to grade it

Costs twenty-one cents and exposes grading bias

A benchmark is only as good as the model you use to grade it

2
Comments 4
9 min read
The model said it read the report. It didn't.

The model said it read the report. It didn't.

2
Comments 1
8 min read