Hey builders
RAGBench just hit Hugging Face.
It's an open evaluation framework for RAG systems, built for developers, with the goal of making RAG evaluation more practical and reproducible.
Instead of guessing whether your chunking strategy actually works, test it. Compare it. Measure it.
I built RAGBench to make it easier to evaluate different RAG approaches using the same documents, queries, and evaluation metrics.
So far, I've focused on comparing semantic, document-level, and parent-child chunking and measuring how they affect retrieval and answer quality.
The goal is simple:
Stop guessing. Start measuring.
RAGBench is:
• Open source
• Open data
• Reproducible
• Built for community feedback
Try the interactive benchmark:
https://huggingface.co/spaces/Gul55555/ragbench
Contributions and feedback are welcome.
I'm especially interested in:
• What RAG evaluation metrics do you actually use?
• What chunking strategies have worked best for you?
• What problems have you encountered when evaluating RAG?
• What should I add to RAGBench next?
What's broken in your RAG pipeline?
I'd love to hear what you're working on.
Top comments (0)