GraphRAG

Closing the shared-meaning gap: what it actually takes

Last week's piece was about the problem: your data has structure but not meaning, and AI now acts on that gap at machine scale. This one is about the work. What should an organisation build, what skills does a data architect or analyst need, and which techniques, languages and standards should they use, in what order?

Start with a diagnosis, not a purchase

The most common failure in this field happens before any modelling starts. Someone says "we need a semantic layer", a vendor demonstrates one, and the budget goes on a tool aimed at a different layer from the actual problem.

Last week I described five senses of "semantic layer": a shared vocabulary, a canonical domain model, the analytics or metrics layer, the operational or integration layer, and formal ontology. The first skill is to tell them apart under pressure. When two dashboards disagree, is that a vocabulary problem (people mean different things), an identity problem (the same customer is four records), or a derivation problem (the same result calculated twice)? Each has a different fix and a different owner.

Two techniques make the diagnosis concrete.

  • Map the domain as a graph. Sketch the ten to twenty things that matter as nodes and their relationships as named edges. Then mark where the current models let you down: meaning that lives only in a column label, identity that is local to one system, a missing row treated as "false", and schemas that are siloed or brittle

  • Write competency questions. Before modelling, write down the questions the model must be able to answer, in business language, and agree them with the people who will ask them (Grüninger & Fox, 1995). They set the scope, they are your acceptance tests, and they stop the model growing to cover everything

None of this needs a tool. It needs someone who can facilitate a room full of people who each think their definition is the obvious one.

1. A governed shared vocabulary

This is the cheapest semantics that pays, and the step most often skipped because it looks too simple. Most "data quality" disputes are vocabulary disputes in a technical costume.

A governed vocabulary means every business term has a preferred label, a written definition, a named owner, its synonyms, and its place in a hierarchy. The skill is writing definitions that settle arguments ("a customer is a natural or legal person who holds at least one live product") rather than restating the word, and then running a change process that people actually use.

  • SKOS — the Simple Knowledge Organization System (W3C Recommendation, 2009) is the standard way to publish it: concept schemes, prefLabel and altLabel, broaderand narrower, and mapping properties (exactMatch, closeMatch) for aligning two vocabularies that grew up apart.

  • ISO 25964 covers thesaurus construction and interoperability; ISO/IEC 11179 covers metadata registries and how to name and define a data element properly.

  • Tools range from a well-governed glossary in your data catalogue to dedicated vocabulary managers such as VocBench, PoolParty or TopBraid EDG.

One confusion is worth dealing with early: SKOS broader is not subclass of. "Espresso machines" can sensibly sit under "Coffee" in a navigation hierarchy, but a machine is not a kind of coffee. Vocabularies organise terms; ontologies make logical claims about things. Mixing them up causes trouble further on.

2. Thinking in graphs

This is the skill shift, and in my experience it is harder for excellent data modellers than for beginners, because their relational instincts are so practiced.

Entity-relationship modelling, UML and Information Engineering remain a solid foundation, and nearly all of it carries over. What changes is a small set of assumptions:

  • Relationships are first-class. A relationship can carry its own properties, be queried directly and be followed across many hops

  • Open world, not closed. In SQL a missing row means "no". In semantic systems a missing fact means "unknown". Queries, validation and reasoning all behave differently as a result

  • No unique-name assumption. Two identifiers might denote the same thing unless you say otherwise. That is exactly the situation integration puts you in

The labelled property graph is the most accessible way to build this skill. It is close enough to relational thinking that a modeller is productive within a day, and it answers the questions relational databases handle worst, such as "every customer who bought a product made from this contaminated lot". Learn a repeatable recipe for translating an ER model into a graph, and the patterns that separate a good graph from a relational schema with arrows on it: model facts as relationships rather than string properties; promote a property to a node when you need to query or relate it; use intermediate nodes for facts involving more than two things.

The languages are Cypher (openCypher, as in Neo4j and others) and now GQL — the Graph Query Language (ISO/IEC 39075:2024), the first new ISO database query language since SQL. SQL/PGQ (part of ISO/IEC 9075:2023) brings graph pattern matching into SQL itself.

3. A canonical model on global identity

A property graph fixes connection. It does not, on its own, fix meaning across systems, because its node identifiers are still local to one database. The step that turns a data model into an interoperability layer is global identity.

  • RDF — the Resource Description Framework reduces everything to the triple (subject, predicate, object) and names things with IRIs — Internationalized Resource Identifiers. Two systems that use the same IRI are making the same claim about the same thing, and their graphs merge without any mapping code. RDF 1.1 (2014) is the current Recommendation; RDF 1.2 reached Candidate Recommendation in April 2026.

  • RDFS — RDF Schema adds classes, properties and hierarchy. The classic trap: rdfs:domain and rdfs:range infer types; they do not constrain data. A modeller who expects them to behave like foreign keys will be surprised

  • Reuse before you invent. schema.org, FIBO (the Financial Industry Business Ontology), SNOMED CT, QUDT for units, DCAT for datasets and PROV-O for provenance already cover much of what enterprises model from scratch. Evaluating a standard is its own skill: is it good as a vocabulary, as logic, or both?

  • Lift the estate you already have. R2RML (W3C, 2012) and RML map relational and file-based sources into the canonical model declaratively, without rewriting the sources

  • SPARQL queries the result and, with SERVICE, runs one query across graphs held by different organisations without copying the data

The architectural pay-off is a hub. n systems integrated point to point need up to n(n−1)/2mappings, each with its own quiet assumptions. With a canonical model, each system maps once, to the hub.

Should it be RDF or a property graph? Increasingly the answer is both, with a bridge between them: property graphs for connected operational workloads, RDF where identity, reuse and meaning have to travel. The skill is choosing per use case rather than per ideology.

4. Formal meaning, in proportion

OWL 2 — the Web Ontology Language (W3C, 2012) lets you say what terms necessarily mean, in a form a reasoner can check and draw conclusions from. It is powerful and, as I argued last week, often the most over-applied technique on the map.

For most practitioners the skill needed is reading, not engineering: understand classes, properties and restrictions; know the difference between a primitive class (necessary conditions only) and a defined class (necessary and sufficient, so a reasoner can classify into it); and predict what a reasoner will and will not infer

Visual notations help analysts here. G-OWL (Héon & Paquette) is a complete visual syntax for OWL 2, so a diagram round-trips to Turtle (RDF serialisation) and back, and VOWL offers a lighter-weight view. Both let a domain expert check a formal model as easily as an ER diagram.

The warning from last week still applies. A language model will draft you an ontology in minutes. Without competency questions, analysis and validation, you get a vibe ontology: fluent, plausible and unverifiable. We can use LLMs to do the grunt work, but we need to validate and then verify.

5. Governed metrics: the analytics semantic layer

Everyone agrees what a customer is, but two dashboards still disagree on the count, because the calculation lives in two places. This is metric drift, and it has a boardroom cost.

The fix is to treat a metric as a contract: defined once, with its entities, dimensions, measures, time grain, owner and version, and then served headlessly, so the dashboard, the notebook and the agent all call the same definition instead of reimplementing it. The tools include the dbt Semantic Layer (MetricFlow), Cube, AtScale and LookML.

The link to the earlier steps is what makes this work. A metric's dimensions should come from the governed vocabulary, and its grain from the canonical model. A metrics layer over four unreconciled customer records computes flawlessly and is still wrong.

6. Data products, contracts and governance

Meaning does not stay shared on its own. In a decentralised estate it needs structure:

  • Data products with semantic contracts. Data mesh's four principles (domain ownership, data as a product, self-serve platform, federated governance) depend on semantics as their connective tissue. A product's "order" should be the canonical model's Order, with its real relationships intact

  • Naming, versioning and stewardship. Conventions people can follow; deprecation rather than deletion; and distinct roles for who defines a term, who owns the data and who approves change

  • Lineage and provenance. Metric-level impact analysis ("what breaks if we change active customer?"), recorded in a standard way with PROV-O — the PROV Ontology

  • Validation at the boundary. The model can stay open-world while data entering a product is checked closed-world. SHACL — the Shapes Constraint Language (W3C, 2017; 1.2 in draft) is the standard for that. OWL says what things are; SHACL says what a valid record looks like

The human side matters as much as the apparatus. A lightweight change process that people use beats a thorough one they avoid.

7. Grounding AI

An agent working over raw tables has no stable notion of an entity, invents joins, and delivers wrong answers with full confidence. Everything above gives it something firmer to stand on: entities with global identity, governed definitions it can call, and paths it can follow and explain.

Practitioners need to understand the retrieval options well enough to choose between them. RAG retrieves text chunks by similarity. GraphRAG retrieves over a knowledge graph, so the answer follows real relationships. KAG (knowledge-augmented generation) goes further, reasoning over the graph's structure and logic. A graph earns its place over a vector index when the question depends on relationships, identity or rules rather than on finding similar text.

Shared meaning is the prerequisite for reliable AI, not the polish on it.

Putting it together: a minimum viable kit

An organisation does not need all of this at once. It needs a small, coherent starting point that the next team can build on. In practice that is five artefacts for one domain, each of which feeds another:

  • A governed vocabulary, which names the model's entities and supplies the metric's dimensions

  • A canonical model, which sets the metric's grain and supplies the data product's entities

  • One governed metric, wired to the model

  • One data product with a stated contract

  • A one-page governance note that gives each an owner, a version and a change process

Add the competency questions the kit can now answer, so whoever picks it up knows what it is for and can test it. That is a defensible first rung, not a half-built system. The next rungs are formal rigour on the model, SHACL validation against real data, and agents that act on the graph. They cost more, and each stands on the one below.

The skills, in one place

Taken together, this describes a role many organisations already need but rarely name: a semantic translator between the business and the data and AI teams. The profile is:

  • Diagnosis: placing a problem on the right layer, and saying so to a sponsor

  • Elicitation: competency questions, and definitions that settle arguments

  • Modelling: ER and UML fluency, plus graph thinking and the open-world gear-change

  • Standards literacy: SKOS, RDF, RDFS, SPARQL, Cypher and GQL at working level; OWL and SHACL at reading level

  • Analytics: governed metrics and a headless semantic layer

  • Governance: ownership, versioning, lineage and data contracts

  • AI literacy: why grounding matters and which retrieval pattern fits which question

Most of it needs no formal logic at all, and it returns most of the value.

How we're approaching it

This is the ground we built our training pathway around. It has two courses, under one line: Build Shared Meaning Today to Engineer Reliable Agentic Automation Tomorrow.

Course 1 — Building Shared Meaning is for analysts, BI practitioners, data modellers and domain experts, and needs no programming or logic background. Over four and a half days it works through capabilities 0 to 6 above at hands-on depth and 7 at awareness level: diagnosis and domain mapping; property graphs with Cypher and GQL; a governed SKOS vocabulary; RDF, RDFS, reading OWL through G-OWL, and SPARQL; governed metrics in dbt and Cube; data products, governance and lineage; and the connection to AI. Each day produces one artefact of the minimum viable kit, and participants leave with the whole kit built for a domain of their own.

Full details here: https://www.inspired.org/course-1-build-shared-meaning

Course 2 — Engineering Reliable Agentic Automation takes the upper rungs: formal rigour, SHACL validation, the operational layer and grounded agents. More on that in the next post.

References and standards

  1. W3C, SKOS Simple Knowledge Organization System Reference, Recommendation, 2009 — https://www.w3.org/TR/skos-reference/

  2. ISO 25964, Thesauri and interoperability with other vocabularies; ISO/IEC 11179, Metadata registries

  3. ISO/IEC 39075:2024, Information technology — Database languages — GQL; ISO/IEC 9075-16:2023, SQL/PGQ

  4. W3C, RDF 1.1 Concepts and Abstract Syntax, Recommendation, 2014 — https://www.w3.org/TR/rdf11-concepts/ · RDF 1.2 Concepts and Abstract Data Model, Candidate Recommendation, April 2026 — https://www.w3.org/TR/rdf12-concepts/

  5. W3C, R2RML: RDB to RDF Mapping Language, Recommendation, 2012 — https://www.w3.org/TR/r2rml/

  6. W3C, SPARQL 1.1 Query Language and Federated Query, Recommendations, 2013 — https://www.w3.org/TR/sparql11-query/ (SPARQL 1.2 in Working Draft)

  7. W3C, OWL 2 Web Ontology Language Primer, 2nd edition, 2012 — https://www.w3.org/TR/owl2-primer/

  8. Héon, M. & Paquette, G., "G-OWL: A Complete Visual Syntax for OWL 2 Ontology Modeling and Communication", Semantic Web journal — semantic-web-journal.net

  9. W3C, Shapes Constraint Language (SHACL), Recommendation, 2017 — https://www.w3.org/TR/shacl/ (SHACL 1.2 in Working Draft)

  10. W3C, PROV-O: The PROV Ontology, Recommendation, 2013 — https://www.w3.org/TR/prov-o/

  11. dbt Labs, dbt Semantic Layer / MetricFlow — docs.getdbt.com · Cube — cube.dev/docs

  12. Grüninger, M. & Fox, M. (1995), "Methodology for the Design and Evaluation of Ontologies", IJCAI Workshop on Basic Ontological Issues in Knowledge Sharing.

  13. Dehghani, Z. (2022), Data Mesh: Delivering Data-Driven Value at Scale, O'Reilly.

  14. Robinson, I., Webber, J. & Eifrem, E. (2015), Graph Databases, 2nd edition, O'Reilly.

  15. Allemang, D., Hendler, J. & Gandon, F. (2020), Semantic Web for the Working Ontologist, 3rd edition, ACM Books.

  16. Edge, D. et al. (2024), "From Local to Global: A Graph RAG Approach to Query-Focused Summarization" — arXiv:2404.16130

  17. Liang, L. et al. (2024), "KAG: Boosting LLMs in Professional Domains via Knowledge Augmented Generation" — arXiv:2409.13731