DEV Community

Rhesis.AI profile picture

Rhesis.AI

Find out where your AI agents fail. Open-source testing for LLM & agentic apps.

Location Germany Joined Joined on  Personal website https://rhesis.ai/ twitter website
We tested our own healthcare agent. It missed five emergencies out of twenty.

We tested our own healthcare agent. It missed five emergencies out of twenty.

Comments
11 min read
Building a Test Harness for Healthcare Conversational AI

Building a Test Harness for Healthcare Conversational AI

Comments
19 min read
Testing conversational AI for healthcare: why it's different

Testing conversational AI for healthcare: why it's different

Comments
14 min read
Watching Isn't Testing: The Case for Two-Way Connections

Watching Isn't Testing: The Case for Two-Way Connections

1
Comments
8 min read
Rhesis got a new UI & a new way to work | Release v0.9

Rhesis got a new UI & a new way to work | Release v0.9

1
Comments
1 min read
Testing LLM & agentic applications | Rhesis AI Product Demo

Testing LLM & agentic applications | Rhesis AI Product Demo

1
Comments
1 min read
Scoped access, managed secrets, self-healing deploys: our move to Kubernetes

Scoped access, managed secrets, self-healing deploys: our move to Kubernetes

2
Comments
7 min read
Offline vs. online evaluation at the application layer: a practical guide

Offline vs. online evaluation at the application layer: a practical guide

Comments
12 min read
Deploying a Custom LLM in Production: Four Architectures, Only One Works

Deploying a Custom LLM in Production: Four Architectures, Only One Works

1
Comments
13 min read
What EvalOps is and why AI teams can't ship without it

What EvalOps is and why AI teams can't ship without it

Comments
8 min read
7 LLM evaluation & testing tools compared (2026)

7 LLM evaluation & testing tools compared (2026)

Comments
7 min read
7 LLM evaluation & testing tools compared (2026)

7 LLM evaluation & testing tools compared (2026)

Comments
7 min read
loading...