Neo4j's official framework for building a knowledge graph is seven steps long. It reads like a well-behaved playbook: define the use case, pick your sources, design the ontology, model the data, ingest, query, operate. If you were reviewing it on a slide it would look completely reasonable.
Then you sit down on a Friday night to actually build one, and Step 3 quietly reaches out and eats the whole weekend.
Twice.
The first time I built a knowledge graph on Neo4j, the ontology-design step ate two full weekends before I could even start writing the first Cypher CREATE statement. Steps 1, 2, 4, 5, 6, and 7 combined took less time than Step 3 by itself. That is not because the other six steps are trivial. It is because Step 3 is where every unresolved assumption from Steps 1 and 2 comes home to be reconciled, and you cannot fake your way past it.
Below is the 7-step arc as Neo4j teaches it, with what actually happens between the diagrams.
The 7 steps, at a glance
| Step | What Neo4j calls it | What actually happens |
|---|---|---|
| 1 | Define the use case | You pick one query you want fast |
| 2 | Identify data sources | You classify structured / semi-structured / unstructured |
| 3 | Design the ontology | You loop back to Step 1 three times |
| 4 | Model the data | You write Cypher and remember indexes |
| 5 | Ingest and transform | LOAD CSV, APOC, LLM extraction |
| 6 | Build queries and APIs | Cypher becomes your business logic |
| 7 | Operate and extend | Freshness, quality monitoring, schema evolution |
Steps 1 and 2 look easy on paper. Step 3 is the interesting one. Steps 4-7 are where the real technology sits.
Step 1: define the use case (or watch the whole project rot)
The bad version of a use case:
"We want to put all our internal data into a knowledge graph."
The good version:
"A new engineer needs to know, in 30 seconds, which services depend on
/api/usersbefore they change it."
What separates the two versions is not scope but specificity. A specific query pins down three things that would otherwise stay loose for months: which entities you need as nodes, which relationships you need as edges, and which query pattern has to be fast.
The industry has a graveyard full of KG projects that started with "let us model everything." Neo4j's own guidance is blunt about this: if the use case is not concrete, the graph is decoration.
Step 2: identify data sources (structured, semi-structured, unstructured)
The classification is the whole step:
- Structured — CSVs, relational tables, customer masters, product catalogs.
- Semi-structured — JSON API responses, config files, OpenAPI specs.
- Unstructured — PDFs, meeting notes, source code, chat logs.
The reason this classification matters is that it decides how you get into Step 5. Structured data walks into Neo4j through LOAD CSV. Semi-structured comes in via APOC. Unstructured needs an LLM extractor between the source and the graph.
In September 2026 the LLM-extractor side of this has become a solved product problem: Neo4j's own LLM Knowledge Graph Builder requires Neo4j 5.23+ and now supports OpenAI, Gemini, Claude, Llama 3, Amazon Nova, Diffbot, and Qwen as the extraction model. The neo4j-graphrag-python SimpleKGPipeline turns on structured-output extraction when the LLM supports it, which cuts the "the model returned invalid JSON" retries you would otherwise write yourself. If you are starting today, you build the extractor last and use these tools.
Step 3: ontology design (the one that ate the weekends)
This is where the trouble happens. On paper, an ontology looks like a tiny document:
Node labels:
- Service (name, version, team)
- API (path, method, status)
- Developer (name, email, team)
- Repository (name, url, language)
Edge types:
- EXPOSES: Service -> API
- CALLS: API -> API
- MAINTAINS: Developer -> Service
- HOSTED_IN: Service -> Repository
That is maybe 15 lines. How did that eat two weekends?
Because writing those 15 lines forces you to answer questions Step 1 quietly deferred. Is a deployed instance of a service a Service or a Deployment? Is a Repository really one node, or one per branch? Does a MAINTAINS edge point from the developer to the service or the other way around, and does the direction change what queries become expensive?
The moment you start writing Cypher against a draft ontology, you find edges that need properties, node labels that need to split in two, and relationships whose direction you had wrong. You go back to Step 1, reread the use case, and change Step 3. Then you go back to Step 3 and change it again. The rough pattern I ended up with: Step 3 makes you loop back to Step 1 about three times before it settles.
The lesson I wish I had absorbed earlier:
- Do not design the ontology in isolation. Sketch it in the same file as three sample Cypher queries you actually plan to run. If any of the three cannot be written cleanly, the ontology is wrong, not the query.
- Start small. Two node labels and one edge type is a legitimate first iteration. Expanding is easy; unpicking is a weekend.
- Read edges as sentences. "Service EXPOSES API." "Developer MAINTAINS Service." If you cannot say the edge out loud without it sounding weird, the direction is off.
I would also, in retrospect, have used Neo4j's Graph Consolidation feature earlier. It rolls a sprawling schema into fewer, more meaningful node labels and relationship types automatically. On a first attempt this is exactly the kind of judgement I did not yet have.
Step 4: data modelling in Cypher
Once the ontology settles, Step 4 is mechanical:
CREATE (:Service {name: "UserAPI", version: "2.1", team: "Platform"});
CREATE (:Service {name: "AuthService", version: "1.5", team: "Security"});
CREATE (:API {path: "/api/users", method: "GET", status: "active"});
MATCH (s:Service {name: "UserAPI"}),
(a:API {path: "/api/users"})
CREATE (s)-[:EXPOSES {since: "2024-01-15"}]->(a);
CREATE INDEX FOR (s:Service) ON (s.name);
CREATE INDEX FOR (a:API) ON (a.path);
The single most common mistake at this step: forgetting the indexes. A KG with millions of nodes and no index on the query's start-node property will take seconds per query, and you will spend a week thinking your ontology is wrong when the real problem is a CREATE INDEX line you never wrote.
Step 5: ingestion and transformation
For structured data, LOAD CSV WITH HEADERS is the whole story:
LOAD CSV WITH HEADERS FROM 'file:///services.csv' AS row
CREATE (:Service {
name: row.name,
version: row.version,
team: row.team
});
For unstructured data, the modern move is to hand the LLM a schema-constrained extraction prompt rather than a free-form one. SimpleKGPipeline in neo4j-graphrag-python does this for you when the underlying model supports structured output; you get JSON that conforms to your ontology rather than JSON that hopes to. The retries you save at this step add up.
Step 6: queries and APIs
This is where the graph starts paying rent. The three query shapes I ended up using most:
Blast radius — what breaks if I change this API?
MATCH (target:API {path: "/api/users"})<-[:CALLS]-(caller:API)<-[:EXPOSES]-(s:Service)
RETURN s.name AS affected_service,
caller.path AS via_api;
Shortest path between two services.
MATCH path = shortestPath(
(a:Service {name: "UserAPI"})-[*]-(b:Service {name: "PaymentService"})
)
RETURN path;
Community detection — which services cluster together?
CALL gds.louvain.stream('service-graph')
YIELD nodeId, communityId
RETURN gds.util.asNode(nodeId).name AS service, communityId
ORDER BY communityId;
Cypher is where a knowledge graph stops being a nice diagram and starts being business logic. Those three shapes cover most of what you actually reach for once the graph is live.
Step 7: operate and extend
The step nobody wants to write about, because it does not screenshot well.
- Freshness. The source data changes; your graph does not, unless you have a sync pipeline.
- Quality. Orphan nodes and duplicate entities creep in. Schedule a job to find and merge them.
- Schema evolution. New use cases mean new node labels and new edges. Version your ontology.
- Access control. Not every team should see every subgraph.
Neo4j AuraDB shoulders most of the infrastructure side of this, which is why a lot of small teams end up on Aura even if they started on a self-hosted instance. The trade is real: less flexibility, less ops.
What I would do differently
If I were starting over today with the same use case:
- Write the three killer queries first, before the ontology. The queries constrain the ontology far better than the ontology constrains the queries.
- Cap the first ontology at two node types and one edge type. Expand only when a query forces the expansion.
-
Use
SimpleKGPipeline(or the LLM Knowledge Graph Builder if you want a UI) instead of hand-writing extraction prompts. The structured-output layer alone saves days. - Do not skip the indexes. Ever.
- Budget for Step 3 to take longer than Steps 1, 2, 4, 5 combined. If it does not, either your use case is unusually narrow or you are about to find out you missed something.
The 7-step arc is honest. It is the shape of the effort inside each step that the diagram does not capture. Once you know Step 3 is going to be the weekend eater, you can plan around it — you sketch queries alongside the ontology, you keep the first pass small, and you accept that "the graph is designed" is a claim you cannot make until you have written Cypher against it.
If you want the longer version, my Zenn book Knowledge Graph Practical Guide walks the full arc, and Chapter 4 is exactly this 7-step build — including the ontology-loop story that turned my weekends into a lesson.
Top comments (0)