Review
So I finally sat down with ODKE+, Apple’s paper on how they keep their knowledge graph fresh. They read the web, pull facts out in the shape of an ontology, check them, and merge them into the graph. I didn’t stop at reading it though. I built the whole thing as an open-source package, openodke (so yes, I have skin in this game, keep that in mind), and then tested it on public data against the two libraries most people would actually reach for.
Let me start with the honest bit, and then you can decide whether the rest is worth your time.
To me it felt less like a technical paper and more like an industry record of what they’ve achieved for their own internal use case. I highly doubt this paper is going to help anyone do significant, replicable work for their own use cases when it comes to building knowledge graphs from text.
That sounds harsher than I mean it. Stay with me.
What the paper actually says
Strip away the diagrams and there are three ideas in it.
- Ontology-guided extraction. For every entity type the model is handed a ranked slice of the ontology, the predicates that matter most for that type, and it has to answer in the ontology’s terms p. 4.
- A grounder. Every extracted fact goes to a second, lighter model, which gets asked one question: does the evidence support this? Yes stays, no goes p. 5. The exact prompt is in the appendix p. 8.
- Corroboration and scoring. The same fact showing up across many sources gets merged and scored: how many sources agree, which extractor found it, how confident it was p. 5.
And the numbers they report are big. 19 million facts at 98.8% precision. Grounding cut hallucinated extractions by 35%. Corroboration took precision from 91% to 98.8% p. 6.
Now my problem is this. Most of what makes those numbers work isn’t in the paper:
- their own knowledge graph, used to enrich the predicates;
- labelled audits that train the scorer;
- rules written for Wikipedia;
- extraction aimed at an entity they already know;
- the whole web repeating the same fact a hundred times over, so corroboration has something to chew on.
Let me draw a parallel. It’s like a restaurant publishing its recipe and keeping the kitchen, the suppliers and the chef. You can follow that recipe to the letter and still not get the dish.
So what did I find?
I ran three systems on two public benchmarks, with the same model, the same documents and the same schema: openodke, LangChain’s LLMGraphTransformer and Neo4j’s own extractor. Then I put openodke’s grounder on top of everyone’s output. (The full setup is further down, if you’re into that sort of thing.)
Four things stood out.
Grounding works. Just not 35% works. On whole documents the grounder knocked out 13 to 19% of the wrong facts. On single sentences it barely moved anything, and sometimes it threw out facts the benchmark said were correct.
Corroboration, the thing behind the paper’s biggest jump, couldn’t even be tested. Every fact in these datasets appears in exactly one place. Nothing to corroborate. Across six runs it changed one fact. ONE.
Then the ontology guidance. It keeps the output in shape, sure, but so does everything else once you hand it the schema strictly. All three systems hit 100% conformance. The ranked slice actually hurt when it hid relations from the model: showing openodke the whole ontology bumped its recall by 4.5 points.
And my own extractor (yes, mine) turned out precise, not thorough. It was the most precise and the cheapest on documents, but it found about half the facts LangChain did.
My verdict
I think this paper has some good ideas, but overall the pipeline is very generic. There are a few good ideas we can definitely take from it. The pipeline itself, though, will be very hard to replicate for use cases that aren’t very close to what Apple was testing. It will definitely fall short in use cases that are very specific to a domain, let’s say sales or cybersecurity, where documents are more contrived, entities are more closed and the possibility of collisions is very high. So in domains like B2B sales and B2C, I don’t think this would work very well, and it definitely won’t scale.
Now I know that last part is a judgment call. I didn’t measure it. More on that at the very end.
The good ideas aren’t limited to this paper either, and the big one is the grounder. That’s probably what separates this paper from its predecessors. The grounder, along with the corroborator that ranks the sources, creates a certain trust value in the post-LLM era of knowledge extraction. It’s a good directional framework for anyone deciding how to use LLMs to build that ingestion pipeline, or at least the extraction and corroboration part of it.
And the experiment backs that up. The grounder doesn’t need the rest of ODKE+. It doesn’t need to be tied to ODKE at all, since it’s a paradigm. It can sit on its own in whatever Text2KG pipeline you build, right after the extraction. I ran it, unchanged, on LangChain’s and Neo4j’s output, and LangChain plus the grounder gave the best balance of anything I tested: 60% precision at 24% recall.
One more thing, so nobody walks away reading this as “the paper is wrong”. I don’t think it’s a highly technical, research-oriented paper, and I don’t think they even claim the results are replicable. They found an approach that worked for them and published what they found. We couldn’t corroborate it on any open-source dataset. That doesn’t mean the paper is wrong, and it doesn’t mean our experiment negates it. It just means this isn’t something you can reproduce as an experiment, at any scale, at least for now. And honestly, I don’t think that was ever the point of the paper.
If you want to try this at your scale
At the end of the day, if you’re a small team and you want ODKE+-style results, this is what I’d actually do.
- Keep whatever extractor you already have, and put a grounder right after it. A small model answering true or false per fact against the whole document. It cost me between a tenth and a fifth of what extraction cost.
- If your ontology is not that big, skip the ranking of snippets and pass the complete ontology to the extractor. Ranking only earns its place once the ontology stops fitting in the prompt.
- Don’t count on corroboration unless the same facts genuinely show up across many of your documents. Count that first.
- In a closed domain, be really careful merging entities. A link you can undo beats a merge you can’t.
- Measure on your own labels. Public benchmarks told me the shape of the trade-offs; your data tells you the actual numbers.
The experiment
Hypotheses
The paper makes three claims that can be checked on public data. I added a fourth.
- Grounding cuts hallucinated extractions by about 35% p. 6.
- Corroboration raises precision from 91% to 98.8% p. 6.
- Ontology-guided extraction keeps the output schema-aligned, and holds its own against common libraries on the same model p. 4.
- Mine: the grounder is separable. It improves any extractor’s output, not only the paper’s.
The headline number, 19 million facts at 98.8% precision, uses Apple’s private graph and audits. You can’t test it in public.
Datasets
| Text2KGBench (Wikidata-TekGen) | Re-DocRED | |
|---|---|---|
| What it is | Single sentences aligned to the Wikidata facts they state; one small ontology per domain | Wikipedia introduction paragraphs annotated for document-level relations |
| Used | 10 ontologies × 20 test sentences = 200 sentences, 319 gold facts | 50 test documents (median 164 words), 1,747 gold facts |
| Schema | 4 to 15 relations per ontology | 96 relations; 52 to 65 per entity type |
| Why | The paper’s own setting, cut down to sentences | The paper works on whole pages. Here 53% of the facts span two sentences, which is the hard case |
Metrics
Precision is the share of a system’s facts that match gold. Recall is the share of gold facts the system found. F1 combines the two into one number.
Ontology conformance is the share of facts whose relation is in the schema. Hallucinated facts follow Text2KGBench’s own definition: the subject or object is not in the sentence, or the relation is not in the ontology.
For the grounder, I count wrong facts (facts that match no gold) before and after it runs. Cost is the provider’s list price for the tokens used.
Matching: Text2KGBench compares facts after lowercasing and removing spaces. Re-DocRED accepts any of an entity’s gold names.
Setup
- Claude Sonnet 5.5 extracts for all three systems, with 16,000 output tokens of room.
- Claude Haiku 4.5 grounds, in the paper’s own mode. It reads the whole document and answers true or false for each fact. Only the true facts stay.
- Every system gets the same documents, whole, and the same schema, with every relation shown.
- Competitor facts outside the schema’s patterns are dropped, which is what LangChain’s strict mode does anyway.
- No value normalisation. “13 March 1963” stays “13 March 1963”, because gold writes it that way.
- Every system calls its model through LiteLLM, so the harness runs on any provider.
Candidates
| System | What it is |
|---|---|
| openodke | My open-source implementation of ODKE+: ranked ontology slices, an exact quote for each fact, then a grounder, corroborator, scorer and validator |
LangChain LLMGraphTransformer | The extractor most LangChain users reach for, in strict schema mode |
neo4j-graphrag LLMEntityRelationExtractor | Neo4j’s own knowledge-graph-builder extractor, with no extra types allowed |
Microsoft GraphRAG isn’t here. It writes free-text descriptions with no ontology, so you can’t score it against a dataset’s relations.
Runs
Each system runs on both datasets in three configurations:
- extraction alone;
- plus grounding;
- plus corroboration.
The competitors’ facts go through openodke’s grounder and corroborator unchanged.
A confession: the first run wasn’t fair. On Re-DocRED, openodke saw at most 25 relations per entity type, while the other two saw all of them. I fixed the bench and ran openodke again. The numbers below are from that rerun.
All of it together cost under $20.
Results
| Each system’s own facts | openodke | LangChain LLMGraphTransformer | neo4j-graphrag |
|---|---|---|---|
| Text2KGBench: precision / recall / F1 | 46.5 / 41.1 / 42.4 | 50.1 / 47.2 / 47.3 | 46.0 / 42.0 / 42.7 |
| Text2KGBench: hallucinated facts | 4.5% | 3.2% | 8.4% |
| Re-DocRED: precision / recall / F1 | 67.9 / 14.4 / 23.8 | 56.2 / 25.1 / 34.7 | 49.2 / 18.5 / 26.9 |
| Re-DocRED: facts extracted | 371 | 781 | 657 |
| Re-DocRED: extraction cost | $1.30 | $2.64 | $2.15 |
What the grounder did to each system’s facts:
| openodke | LangChain | neo4j-graphrag | |
|---|---|---|---|
| Re-DocRED: wrong facts, before → after | 119 → 102 (−14%) | 342 → 278 (−19%) | 334 → 291 (−13%) |
| Re-DocRED: precision after grounding | 70.9% | 60.2% | 52.0% |
| Text2KGBench: change in wrong facts (approximate) | −1% | −4% | −6% |
| Grounding cost, both datasets | $0.32 | $0.60 | $0.51 |
The Text2KGBench changes are approximate, because its precision is averaged over the ten ontologies. Corroboration changed one fact across all six runs.
Interpretation
The systems trade precision against recall, and which one “wins” depends on what you care about. On documents, openodke was the most precise and the cheapest, but it found about half of what LangChain found. On single sentences, LangChain led on everything, and openodke sat level with Neo4j.
Grounding behaves differently on documents and sentences. On documents it mostly removed wrong facts. On sentences it removed more facts that gold called right than ones gold called wrong. Some of that is gold’s fault: Text2KGBench is distantly supervised, so some of its “right” facts aren’t actually stated in the sentence.
The grounder also does the most for the extractors that don’t quote their evidence. It refused about four times as many of LangChain’s facts as openodke’s.
Corroboration had nothing to work with. One source per fact.
And schema guidance didn’t separate anyone. Everyone hit 100% conformance with a strict schema. What hurt was hiding relations: the whole ontology gave openodke 4.5 more points of recall.
Conclusions, and what they back in the review
| Conclusion | Backs this in the review |
|---|---|
| The grounder works on other extractors’ output; LangChain plus the grounder gave the best balance (60% precision, 24% recall) | The grounder is the idea worth taking, and it isn’t tied to ODKE |
| Grounding cut wrong facts by 13–19% on documents, not 35% | “Grounding works. Just not 35% works.” |
| The paper’s biggest lever, corroboration, needs many sources per fact, which small corpora don’t have | Very hard to replicate for use cases not close to Apple’s |
| The headline numbers couldn’t be reproduced on public data | Not something you can reproduce as an experiment, at least for now |
| The whole ontology beat a ranked slice at this size | Pass the complete ontology if it’s not that big |
What this cannot tell us
This is just my experiment, and it has edges.
It can’t tell you whether Apple’s system hits its numbers in Apple’s setting. I tested an implementation of the paper, on different data, with one source per fact.
It can’t tell you what corroboration really adds. That needs a dataset where the same facts repeat across many documents, and that’s the next test.
It can’t back my closed-domain claim. That ODKE+ breaks where entities are closed and names collide (sales, cybersecurity) is my judgment. I didn’t measure any domain data.
It says nothing about large scale: ontologies with hundreds or thousands of relations, long documents, or other model families.
And it can’t separate results that are a few points apart. It’s 200 sentences and 50 documents, the gold is incomplete, and nobody audited anything by hand. Trust the shape, not the decimals.