I've been writing software for years – 10 years now – starting with the blue screen of Turbo C, Android, React and now AI. Assisted AI development is not an issue, but applied AI, LLM integrations, RAG, tuning, etc., seem like a whole different world to me.
I ask myself this question every day, and a lot of times, do I need to learn all of this? Am I afraid of being left behind or being archived like an old library? Anyways, I'm finding time to build small things on AI.
I have recently started to learn the concepts of RAG, embeddings and LLM integrations. I came across GraphRAG as well, which seemed like an awesome idea.
GraphRAG was actually more interesting and surprising for me, trying to understand how link traversal works and how indexing happens.
To really see the process, I built GraphRAG Chats, an open-source React Flow app. You draw a graph by hand, prepare it for search, ask it a question, and watch which nodes and relationships the answer came from.
1. GraphRAG
RAG (retrieval-augmented generation) means the AI looks things up in your data before it answers. Plain RAG finds text chunks that are similar to the question.
GraphRAG adds relationships. Your data lives in a graph database such as Neo4j, as nodes. Neo4j can also store embeddings in a vector index, so one database handles both parts: finding the right starting node by meaning and then following its relationships.
2. Embeddings
What are embeddings? Numbers. A list of a few hundred or a few thousand of them.
What do the numbers mean? They capture the meaning of a piece of text. Texts with a similar meaning get similar numbers, so you can measure how close two meanings are.
Where does the text come from? From your own data. You turn each piece of data into a plain, human-readable sentence first, for example, "Ben is a person. Backend engineer," and send that text to an embedding model.
3. Process Step by Step
Preparing the data (done once):
- Create the graph and define all your relationships properly.
- Save it to a graph database (Neo4j in my case).
- Turn each node into readable text.
- Send that text to an embedding model, and store the embedding on each node in Neo4j's vector index.
Answering a question (done every time):
- Turn the question into an embedding with the same model.
- Use vector search to find the nodes whose embeddings are closest to the question. These are the starting points.
- Traverse the graph from those nodes to collect their relationships and neighbours.
- Send the question plus everything found to an LLM, which writes the answer.
Step 8 is the "generation" in retrieval-augmented generation. The LLM only writes the answer; the graph decides what facts it gets to see.
Most of this turned out to be skills I already had as a developer: modelling data, calling APIs, storing things in a database, and debugging when the result is wrong. The new part is a handful of concepts, and they click much faster when you can see them happen.
I would recommend that other developers as well learn by building and do not hesitate to make mistakes. AI transformation is a difficult time, but let's go through it by keeping up with it.


Top comments (0)