Knowledge Graph Construction From Documentation Content
Extracting structured knowledge from messy docs requires resolving entities, not just reading text.

Documentation holds most of what an organization actually knows, but that knowledge sits in a form no system can query directly, which is the entire reason a construction pipeline needs to exist. A compliance question, a data lineage trace, and a cross-system dependency all span documents that RAG retrieval cannot chain on its own, so none of these live inside one document. They span a handful of files, maybe dozens, and retrieval alone can't stitch them together into an answer, because a retriever pulls back passages, not chains of reasoning. As Anthropic's Claude Cookbook puts it, a pile of unstructured documents and questions that span them means no single document contains the answer.
A knowledge graph reframes the problem instead of trying to brute-force it with better search. Entities become nodes, the relationships between them become typed edges, and what used to require repeated retrieval attempts becomes a traversal: walk from node to node, following the edges that answer the question. That's a different computational shape than search, and it's why the exercise is worth the weeks it takes.
But documentation is a harder starting material than the clean corpora used in most NLP research. None of that is fatal, but all of it degrades naive extraction if the pipeline doesn't account for it. Research on this distinguishes between what retrieval does and what a graph does: retrieval surfaces content, but it doesn't make that content authoritative, and it doesn't make it fit to reason from. Closing that gap is the actual job of the five stages that follow, and each one exists because of a specific way documentation resists being turned into structure.
The five-stage pipeline in outline and the points where builds stall
Five stages take an organization from raw documentation to a graph an agent can query: assess and scope, design the ontology and schema, extract entities and relationships, resolve entities, and validate and activate. Atlan's breakdown, published in mid-2026, attaches rough timelines to each: scoping runs one to two weeks, ontology design two to four, extraction one to three, resolution two to four, and validation one to two.
The academic literature groups this differently, as three broader phases: information extraction, information fusion, and information processing. That framing isn't wrong, but it doesn't match how a data team actually staffs and schedules the work, which is why the five-stage view carries more of this article's weight.
This inverts what most people assume going in. Extraction, the stage that sounds hardest (teaching a machine to read and understand unstructured text) is now the cheap part. Entity resolution and ontology alignment, which sound almost administrative by comparison, are where builds actually stall. That reframe shapes the rest of this piece: extraction gets treated as a genuine LLM-enabled success story, while resolution and alignment get the harder, more honest treatment they deserve. Why would de-duplicating a list of entities be harder than reading and understanding a document well enough to extract facts from it? Atlan (published 2026-07-22) sets out the five-stage breakdown with indicative timelines. The total build time runs 6 to 12 weeks for a single domain, or 2 to 3 months enterprise-wide.
Stage 1: Scoping the graph before touching any documents
The single highest-leverage decision in the whole pipeline happens before anyone touches a document: what is this graph actually for, and how wide does it need to reach. Atlan's guidance calls for inventorying the structured sources (data warehouses, BI tools, core business systems) and the unstructured ones (docs, wikis) against one bounded first domain, and treats drawing that boundary as real work, not a formality to get through before the interesting part starts.
Skipping that step produces a predictable failure mode. Scoping isn't glamorous, but it's the thing that decides which metadata sources feed the graph, who has authority over schema decisions, and what "finished" even looks like, none of which can be sorted out mid-extraction once candidate nodes are already piling up.
There's also a prerequisite that's easy to gloss over: governed metadata access. Reverse-constructing a graph from documentation assumes there's already some substrate of metadata to build from, not a blank page. Without it, extraction produces nodes nobody in the organization can actually vouch for. Scoping needs an owner with real cross-team standing, because that same authority gets called on again at the resolution stage, when someone has to decide if two records describing "the same" system really are the same system. An unowned effort at this point is the single most common thread running through abandoned builds.
What comes out of this stage should be narrow, almost disappointingly so: one bounded domain, one named schema owner, one agreed validation framework such as SHACL. Not a sweeping enterprise ontology. That narrowness is the point, not a compromise.
Stage 2: Designing an ontology that survives contact with real documents
Ontology design is where the semantic decisions actually get made, and the instinct to get it all down on paper before writing a single extraction prompt is understandable. It's also the instinct that reliably backfires. The central design task at this stage is figuring out, for any business term that could plausibly be modeled either way, if it's an entity or a property, and that call needs domain experts sitting at the table, not just engineers reasoning from a whiteboard.
Standards such as RDF and OWL should be referenced selectively, not treated as a mandatory authoring project, and the LLM-TEXT2KG 2026 workshop lists "seamless ontology alignment" as a topic for discussion, signaling this is not a solved problem. These rarely hold up, because documentation naming is inconsistent in ways no one anticipates until they're staring at the third variant of the same field name across three different files. RDF and OWL serve as vocabulary and tooling to draw from, but treating either as a mandatory up-front authoring project tends to produce a beautiful ontology that the actual documents don't fit into. The LLM-TEXT2KG workshop at ESWC 2026 in Dubrovnik lists "seamless ontology alignment" as an open topic for discussion, not a settled technique, showing this isn't a solved problem even in the research community.
One useful pattern for containing the scope of ontology design comes from the Knowledge Card architecture, described in a July 2026 paper by Ferreira at Mondegreen.ai. Rather than modeling everything up front, a Knowledge Card captures one bounded concept and attaches five properties to it: ontology-grounded, provenance-linked, boundary-explicit, expert-validated, and versioned. The card also records the conditions under which its own reasoning stops holding, which is really the ontology boundary problem stated out loud instead of left implicit. That idea resurfaces later, in validation, where it matters even more.
The entity types a pipeline needs to support have to be settled before extraction starts, because extraction prompts get structured around them. The full relationship vocabulary, though, doesn't need to be locked down yet and can be refined as real documents surface edge cases. Teams that blur the line between ontology design and semantic-layer design tend to duplicate work a data catalog is probably already doing somewhere else in the org, and pay for it later.
Stage 3: Entity extraction and the limits of what LLMs changed
This is where the economics genuinely shifted, specifically in what changed rather than just gesturing at "AI got better." The classical pipeline needed three separate components built and maintained, a named-entity recognizer trained on domain-specific data, a relation classifier, and hand-written heuristics for entity resolution. Each of those was its own project, with its own training data requirements and its own maintenance burden.
An LLM collapses all three into a single structured-output prompt run against each document, returning entity type, a short description, and subject-predicate-object triples as a validated schema object, no training data and no regex parsing required. Named entity recognition, relation extraction, and coreference resolution (finding every mention across a pile of documents that points to the same underlying entity) now run as prompts against a general-purpose model instead of as three separately trained classifiers.
Anthropic's Claude Cookbook lays out a two-model pattern, understood as an architecture, not a specific product pitch. A fast, cheap model handles the high-volume work: running schema-constrained extraction across every document in the pile. A more capable model gets reserved for entity resolution and summarization, the tasks that require weighing conflicting evidence across documents rather than just following a schema. That division of labor matters, because it's an admission, baked into the architecture itself, that not all of this work is equally hard.
Which raises the obvious question: if extraction is this much cheaper now, why does the pipeline still take weeks? Because cheaper isn't the same as accurate everywhere. Accuracy drops off noticeably on domain-specific terms such as technical codes, operational abbreviations, and spatiotemporal shorthand that appears in engineering reports and assumes local context the model doesn't have. And there's a hallucination problem sitting underneath all of it. Extraction hallucination rates of roughly 10–25% false triples on typical LLMs mean raw streaming output cannot go directly to the graph (filtering layers are required before any triple reaches storage). That means raw extraction output can't stream straight into the graph. That's not caution for caution's sake, it's a structural requirement of the pipeline.
Documentation makes this worse in a specific way: extraction inherits every inconsistency already present in the source material. Inconsistent naming produces duplicate nodes. Missing definitions produce entities nobody can pin down precisely. Undocumented lineage just leaves gaps the model has no way to fill. The worse the source governance, the harder the next stage gets, and there's no extraction technique clever enough to route around that.
Stage 4: Entity resolution, why de-duplication is where enterprise builds stall
The stage that actually earns the pipeline its reputation for difficulty does so for a specific reason. It's not that de-duplication algorithms are primitive. An LLM can weigh conflicting descriptions across a stack of documents and make a judgment call about whether two mentions point to the same real-world thing, which is why the Cookbook's two-model pattern hands this work to the more capable model rather than the fast, cheap one used for bulk extraction.
So why does this stage still eat two to four weeks even with a capable model doing the reasoning? Because the decision of whether two entities should merge isn't a purely linguistic question, it's a question of domain authority, and most organizations don't have one clear owner for that authority spanning every team the graph touches. Whether they actually should be merged, given how finance and sales each use that account differently, is a business decision, not a language-understanding problem. That is the wall the article describes, formed by finance and sales each using that account differently.
The structural roots of this trace straight back to source documentation. Inconsistent naming across docs creates duplicate nodes that all describe the same thing under different labels. Missing definitions leave entities ambiguous enough that even a careful reviewer can't confidently merge them. Undocumented lineage just leaves holes. Governed metadata, meaning clear definitions, consistent naming conventions, and lineage that's actually tracked, is the foundation that makes resolution tractable at any real scale. Without it, resolution turns into an endless argument dressed up as a technical task.
Resolution should not be treated as something done once and shelved. Documentation keeps changing, new surface-form variants keep appearing, and a graph with no ongoing resolution process slowly drifts away from the reality it was supposed to represent.
Morgan Stanley's approach, presented at KGC 2026 at Cornell Tech, offers a concrete look at what taking this seriously looks like in production. Their "Bronze-Silver-Gold" architecture builds governance and domain-expert validation into every layer before data is allowed to advance to the next one, using OntoViewer, a semantic knowledge modeling tool built with metaphacts, as part of the stack. The point of the architecture isn't the specific tooling, it's the structural commitment: resolution and validation as continuous processes, not a one-time cleanup pass.
Agentic systems are starting to take on some of this burden. Athulya Anil's Agentic GraphRAG work, shown at NODES AI 2026 using Neo4j, demonstrates specialized agents automating schema inference and conflict resolution, cutting down the manual load considerably.
Stage 5: Validation and the knowledge infrastructure that makes a graph usable in production
It is the stage that decides if an agent acting on the graph's output can actually be trusted, not a quality-assurance afterthought tacked on at the end, because a hallucinated triple and a correct one look identical at query time. There's no visual tell, no flag, nothing distinguishing a fabricated relationship from a real one once it's sitting in the graph, so the check has to happen before that point.
SHACL, the Shapes Constraint Language, does a lot of this enforcement work. It checks the extracted graph against the ontology's contract: flagging nodes that violate type constraints, edges that connect entity types that were never supposed to connect, properties marked as required that are simply missing.
Human review at this point is doing something different from the domain-expert work back in Stage 2. Stage 2 experts defined what the schema should capture in the first place. Stage 5 experts check whether what actually got extracted matches that intent.
Provenance tracking is easy to underrate until a graph reaches real scale, at which point it becomes load-bearing. A widely cited 2023 survey on knowledge graph construction found that production-grade pipelines need provenance tracking and repair mechanisms built in as graphs grow, because without a way to trace a false triple back to the document it came from, there's no way to correct it. At that scale, a graph without provenance isn't really fixable, it's just something to eventually rebuild.
The Knowledge Card architecture from earlier in this pipeline turns out to be the natural checklist for this stage: ontology-grounded, provenance-linked, boundary-explicit, expert-validated, versioned. Those five properties are effectively the pass/fail criteria for determining if a subgraph is actually ready to feed an agent, rather than just technically present in the database.
The boundary-explicit property matters more for documentation-derived graphs than it might first appear. A Knowledge Card records the specific conditions under which its own reasoning stops being valid. A graph with no sense of its own staleness isn't more useful for looking complete. It's more dangerous, precisely because it looks finished.
Sources
- Knowledge Graph Construction for AI: Enterprise Data Graph (2026)
- Knowledge graph construction with Claude | Claude Cookbook
- metaphacts at The Knowledge Graph Conference 2026
- LLM-TEXT2KG 2026: 5th International Workshop on LLM-Integrated Knowledge Graph Generation from Text (Text2KG)
- A Survey on Semantic Processing Techniques
- Video: NODES AI 2026 - Agentic GraphRAG: Autonomous Knowledge Graph Construction and Adaptive Retrieval


