I Built My Own Database From Scratch — Meet CleaveDB
I've been building something I've wanted to experiment with for a while: CleaveDB.
CleaveDB is a hybrid relational-document graph database built with Python, Rust, and C++ SIMD, with its own query language, graph relationships, and built-in semantic search.
The idea started with a simple question:
What if relational data, documents, graphs, and semantic search could live together in one database?
What makes CleaveDB different?
CleaveDB has its own query language called CleaveQL.
For example:
POUR INTO users "alice" {
"name": "Alice",
"age": 28
}
You can then query the data:
SCOOP users
It also supports graph-style relationships called Bonds:
LINK "users:alice"
TO "users:bob"
AS MUTUAL "friend"
And you can traverse those relationships:
FIND "friend" OF "users:alice"
There's also semantic search:
FIND products
MEANING "something warm for cold weather"
LIMIT 3
So instead of searching only for exact keywords, CleaveDB can search based on meaning.
Under the hood
CleaveDB isn't built on top of an existing database engine. I'm building the storage layer myself using Rust, with C++ SIMD optimisations for performance-critical operations.
The project also includes:
- B+Tree storage
- WAL and transactions
- Graph relationships
- Semantic/vector search
- Authentication and multi-tenant isolation
- Python bindings
- HTTP, WebSocket and TCP interfaces
It's still very much a work in progress, but building the pieces from scratch has been a really interesting learning experience.
Why build another database?
There are already excellent databases like PostgreSQL, MongoDB, Redis, and Neo4j.
I'm not trying to replace them.
CleaveDB is an experiment to see what happens when you combine relational, document, graph, and semantic capabilities into one system.
If you're interested in databases, Rust, storage engines, graph databases, or just want to see how a database is built from scratch, I'd love to hear your thoughts.
Tribrix23
/
CleaveDB
Beyond NoSQL. A hyper-fast graph and vector database you query in plain English. Built on a unique Python, Rust, and C++ SIMD architecture with built-in AI semantic search and strict multi-tenant isolation.
The polyglot, AVX-512 ready , hybrid relational-document graph database with Transformer attention layers.
CleaveDB 3.9.0 is a ground-up, hybrid Relational Document & Graph database that eliminates the complexity of traditional SQL JOINs, external vector search services, and opaque graph databases. It ships with:
-
Native Graph-Relational Links: Documents are not isolated. They are deeply relational, linked natively through ~10 functional Graph "Bonds" that allow unlimited depth, multi-hop traversals with conversational English, completely eliminating the need for
JOINs. - A high-performance Rust storage engine built entirely from scratch — no SQLite, no RocksDB, no external storage libraries.
- AVX-512 / AVX2 C++ SIMD extensions exist in the Rust engine (vector search runs on C++ AVX-512 extensions).
- A Python interpreter frontend (via PyO3 bindings) that runs the CleaveQL query language.
- Real neural Transformer embeddings for semantic search via a quantized ONNX model, using ~22MB of RAM.
- A Go-based distributed coordinator for multi-shard…
I'm especially interested in criticism and ideas for what I should improve next.
Top comments (1)
Pretty cool and it's open source