semantic layer

Closing the shared-meaning gap: what it actually takes

Last week's piece was about the problem: your data has structure but not meaning, and AI now acts on that gap at machine scale. This one is about the work. What should an organisation build, what skills does a data architect or analyst need, and which techniques, languages and standards should they use, in what order?

Start with a diagnosis, not a purchase

The most common failure in this field happens before any modelling starts. Someone says "we need a semantic layer", a vendor demonstrates one, and the budget goes on a tool aimed at a different layer from the actual problem.

Last week I described five senses of "semantic layer": a shared vocabulary, a canonical domain model, the analytics or metrics layer, the operational or integration layer, and formal ontology. The first skill is to tell them apart under pressure. When two dashboards disagree, is that a vocabulary problem (people mean different things), an identity problem (the same customer is four records), or a derivation problem (the same result calculated twice)? Each has a different fix and a different owner.

Two techniques make the diagnosis concrete.

  • Map the domain as a graph. Sketch the ten to twenty things that matter as nodes and their relationships as named edges. Then mark where the current models let you down: meaning that lives only in a column label, identity that is local to one system, a missing row treated as "false", and schemas that are siloed or brittle

  • Write competency questions. Before modelling, write down the questions the model must be able to answer, in business language, and agree them with the people who will ask them (Grüninger & Fox, 1995). They set the scope, they are your acceptance tests, and they stop the model growing to cover everything

None of this needs a tool. It needs someone who can facilitate a room full of people who each think their definition is the obvious one.

1. A governed shared vocabulary

This is the cheapest semantics that pays, and the step most often skipped because it looks too simple. Most "data quality" disputes are vocabulary disputes in a technical costume.

A governed vocabulary means every business term has a preferred label, a written definition, a named owner, its synonyms, and its place in a hierarchy. The skill is writing definitions that settle arguments ("a customer is a natural or legal person who holds at least one live product") rather than restating the word, and then running a change process that people actually use.

  • SKOS — the Simple Knowledge Organization System (W3C Recommendation, 2009) is the standard way to publish it: concept schemes, prefLabel and altLabel, broaderand narrower, and mapping properties (exactMatch, closeMatch) for aligning two vocabularies that grew up apart.

  • ISO 25964 covers thesaurus construction and interoperability; ISO/IEC 11179 covers metadata registries and how to name and define a data element properly.

  • Tools range from a well-governed glossary in your data catalogue to dedicated vocabulary managers such as VocBench, PoolParty or TopBraid EDG.

One confusion is worth dealing with early: SKOS broader is not subclass of. "Espresso machines" can sensibly sit under "Coffee" in a navigation hierarchy, but a machine is not a kind of coffee. Vocabularies organise terms; ontologies make logical claims about things. Mixing them up causes trouble further on.

2. Thinking in graphs

This is the skill shift, and in my experience it is harder for excellent data modellers than for beginners, because their relational instincts are so practiced.

Entity-relationship modelling, UML and Information Engineering remain a solid foundation, and nearly all of it carries over. What changes is a small set of assumptions:

  • Relationships are first-class. A relationship can carry its own properties, be queried directly and be followed across many hops

  • Open world, not closed. In SQL a missing row means "no". In semantic systems a missing fact means "unknown". Queries, validation and reasoning all behave differently as a result

  • No unique-name assumption. Two identifiers might denote the same thing unless you say otherwise. That is exactly the situation integration puts you in

The labelled property graph is the most accessible way to build this skill. It is close enough to relational thinking that a modeller is productive within a day, and it answers the questions relational databases handle worst, such as "every customer who bought a product made from this contaminated lot". Learn a repeatable recipe for translating an ER model into a graph, and the patterns that separate a good graph from a relational schema with arrows on it: model facts as relationships rather than string properties; promote a property to a node when you need to query or relate it; use intermediate nodes for facts involving more than two things.

The languages are Cypher (openCypher, as in Neo4j and others) and now GQL — the Graph Query Language (ISO/IEC 39075:2024), the first new ISO database query language since SQL. SQL/PGQ (part of ISO/IEC 9075:2023) brings graph pattern matching into SQL itself.

3. A canonical model on global identity

A property graph fixes connection. It does not, on its own, fix meaning across systems, because its node identifiers are still local to one database. The step that turns a data model into an interoperability layer is global identity.

  • RDF — the Resource Description Framework reduces everything to the triple (subject, predicate, object) and names things with IRIs — Internationalized Resource Identifiers. Two systems that use the same IRI are making the same claim about the same thing, and their graphs merge without any mapping code. RDF 1.1 (2014) is the current Recommendation; RDF 1.2 reached Candidate Recommendation in April 2026.

  • RDFS — RDF Schema adds classes, properties and hierarchy. The classic trap: rdfs:domain and rdfs:range infer types; they do not constrain data. A modeller who expects them to behave like foreign keys will be surprised

  • Reuse before you invent. schema.org, FIBO (the Financial Industry Business Ontology), SNOMED CT, QUDT for units, DCAT for datasets and PROV-O for provenance already cover much of what enterprises model from scratch. Evaluating a standard is its own skill: is it good as a vocabulary, as logic, or both?

  • Lift the estate you already have. R2RML (W3C, 2012) and RML map relational and file-based sources into the canonical model declaratively, without rewriting the sources

  • SPARQL queries the result and, with SERVICE, runs one query across graphs held by different organisations without copying the data

The architectural pay-off is a hub. n systems integrated point to point need up to n(n−1)/2mappings, each with its own quiet assumptions. With a canonical model, each system maps once, to the hub.

Should it be RDF or a property graph? Increasingly the answer is both, with a bridge between them: property graphs for connected operational workloads, RDF where identity, reuse and meaning have to travel. The skill is choosing per use case rather than per ideology.

4. Formal meaning, in proportion

OWL 2 — the Web Ontology Language (W3C, 2012) lets you say what terms necessarily mean, in a form a reasoner can check and draw conclusions from. It is powerful and, as I argued last week, often the most over-applied technique on the map.

For most practitioners the skill needed is reading, not engineering: understand classes, properties and restrictions; know the difference between a primitive class (necessary conditions only) and a defined class (necessary and sufficient, so a reasoner can classify into it); and predict what a reasoner will and will not infer

Visual notations help analysts here. G-OWL (Héon & Paquette) is a complete visual syntax for OWL 2, so a diagram round-trips to Turtle (RDF serialisation) and back, and VOWL offers a lighter-weight view. Both let a domain expert check a formal model as easily as an ER diagram.

The warning from last week still applies. A language model will draft you an ontology in minutes. Without competency questions, analysis and validation, you get a vibe ontology: fluent, plausible and unverifiable. We can use LLMs to do the grunt work, but we need to validate and then verify.

5. Governed metrics: the analytics semantic layer

Everyone agrees what a customer is, but two dashboards still disagree on the count, because the calculation lives in two places. This is metric drift, and it has a boardroom cost.

The fix is to treat a metric as a contract: defined once, with its entities, dimensions, measures, time grain, owner and version, and then served headlessly, so the dashboard, the notebook and the agent all call the same definition instead of reimplementing it. The tools include the dbt Semantic Layer (MetricFlow), Cube, AtScale and LookML.

The link to the earlier steps is what makes this work. A metric's dimensions should come from the governed vocabulary, and its grain from the canonical model. A metrics layer over four unreconciled customer records computes flawlessly and is still wrong.

6. Data products, contracts and governance

Meaning does not stay shared on its own. In a decentralised estate it needs structure:

  • Data products with semantic contracts. Data mesh's four principles (domain ownership, data as a product, self-serve platform, federated governance) depend on semantics as their connective tissue. A product's "order" should be the canonical model's Order, with its real relationships intact

  • Naming, versioning and stewardship. Conventions people can follow; deprecation rather than deletion; and distinct roles for who defines a term, who owns the data and who approves change

  • Lineage and provenance. Metric-level impact analysis ("what breaks if we change active customer?"), recorded in a standard way with PROV-O — the PROV Ontology

  • Validation at the boundary. The model can stay open-world while data entering a product is checked closed-world. SHACL — the Shapes Constraint Language (W3C, 2017; 1.2 in draft) is the standard for that. OWL says what things are; SHACL says what a valid record looks like

The human side matters as much as the apparatus. A lightweight change process that people use beats a thorough one they avoid.

7. Grounding AI

An agent working over raw tables has no stable notion of an entity, invents joins, and delivers wrong answers with full confidence. Everything above gives it something firmer to stand on: entities with global identity, governed definitions it can call, and paths it can follow and explain.

Practitioners need to understand the retrieval options well enough to choose between them. RAG retrieves text chunks by similarity. GraphRAG retrieves over a knowledge graph, so the answer follows real relationships. KAG (knowledge-augmented generation) goes further, reasoning over the graph's structure and logic. A graph earns its place over a vector index when the question depends on relationships, identity or rules rather than on finding similar text.

Shared meaning is the prerequisite for reliable AI, not the polish on it.

Putting it together: a minimum viable kit

An organisation does not need all of this at once. It needs a small, coherent starting point that the next team can build on. In practice that is five artefacts for one domain, each of which feeds another:

  • A governed vocabulary, which names the model's entities and supplies the metric's dimensions

  • A canonical model, which sets the metric's grain and supplies the data product's entities

  • One governed metric, wired to the model

  • One data product with a stated contract

  • A one-page governance note that gives each an owner, a version and a change process

Add the competency questions the kit can now answer, so whoever picks it up knows what it is for and can test it. That is a defensible first rung, not a half-built system. The next rungs are formal rigour on the model, SHACL validation against real data, and agents that act on the graph. They cost more, and each stands on the one below.

The skills, in one place

Taken together, this describes a role many organisations already need but rarely name: a semantic translator between the business and the data and AI teams. The profile is:

  • Diagnosis: placing a problem on the right layer, and saying so to a sponsor

  • Elicitation: competency questions, and definitions that settle arguments

  • Modelling: ER and UML fluency, plus graph thinking and the open-world gear-change

  • Standards literacy: SKOS, RDF, RDFS, SPARQL, Cypher and GQL at working level; OWL and SHACL at reading level

  • Analytics: governed metrics and a headless semantic layer

  • Governance: ownership, versioning, lineage and data contracts

  • AI literacy: why grounding matters and which retrieval pattern fits which question

Most of it needs no formal logic at all, and it returns most of the value.

How we're approaching it

This is the ground we built our training pathway around. It has two courses, under one line: Build Shared Meaning Today to Engineer Reliable Agentic Automation Tomorrow.

Course 1 — Building Shared Meaning is for analysts, BI practitioners, data modellers and domain experts, and needs no programming or logic background. Over four and a half days it works through capabilities 0 to 6 above at hands-on depth and 7 at awareness level: diagnosis and domain mapping; property graphs with Cypher and GQL; a governed SKOS vocabulary; RDF, RDFS, reading OWL through G-OWL, and SPARQL; governed metrics in dbt and Cube; data products, governance and lineage; and the connection to AI. Each day produces one artefact of the minimum viable kit, and participants leave with the whole kit built for a domain of their own.

Full details here: https://www.inspired.org/course-1-build-shared-meaning

Course 2 — Engineering Reliable Agentic Automation takes the upper rungs: formal rigour, SHACL validation, the operational layer and grounded agents. More on that in the next post.

References and standards

  1. W3C, SKOS Simple Knowledge Organization System Reference, Recommendation, 2009 — https://www.w3.org/TR/skos-reference/

  2. ISO 25964, Thesauri and interoperability with other vocabularies; ISO/IEC 11179, Metadata registries

  3. ISO/IEC 39075:2024, Information technology — Database languages — GQL; ISO/IEC 9075-16:2023, SQL/PGQ

  4. W3C, RDF 1.1 Concepts and Abstract Syntax, Recommendation, 2014 — https://www.w3.org/TR/rdf11-concepts/ · RDF 1.2 Concepts and Abstract Data Model, Candidate Recommendation, April 2026 — https://www.w3.org/TR/rdf12-concepts/

  5. W3C, R2RML: RDB to RDF Mapping Language, Recommendation, 2012 — https://www.w3.org/TR/r2rml/

  6. W3C, SPARQL 1.1 Query Language and Federated Query, Recommendations, 2013 — https://www.w3.org/TR/sparql11-query/ (SPARQL 1.2 in Working Draft)

  7. W3C, OWL 2 Web Ontology Language Primer, 2nd edition, 2012 — https://www.w3.org/TR/owl2-primer/

  8. Héon, M. & Paquette, G., "G-OWL: A Complete Visual Syntax for OWL 2 Ontology Modeling and Communication", Semantic Web journal — semantic-web-journal.net

  9. W3C, Shapes Constraint Language (SHACL), Recommendation, 2017 — https://www.w3.org/TR/shacl/ (SHACL 1.2 in Working Draft)

  10. W3C, PROV-O: The PROV Ontology, Recommendation, 2013 — https://www.w3.org/TR/prov-o/

  11. dbt Labs, dbt Semantic Layer / MetricFlow — docs.getdbt.com · Cube — cube.dev/docs

  12. Grüninger, M. & Fox, M. (1995), "Methodology for the Design and Evaluation of Ontologies", IJCAI Workshop on Basic Ontological Issues in Knowledge Sharing.

  13. Dehghani, Z. (2022), Data Mesh: Delivering Data-Driven Value at Scale, O'Reilly.

  14. Robinson, I., Webber, J. & Eifrem, E. (2015), Graph Databases, 2nd edition, O'Reilly.

  15. Allemang, D., Hendler, J. & Gandon, F. (2020), Semantic Web for the Working Ontologist, 3rd edition, ACM Books.

  16. Edge, D. et al. (2024), "From Local to Global: A Graph RAG Approach to Query-Focused Summarization" — arXiv:2404.16130

  17. Liang, L. et al. (2024), "KAG: Boosting LLMs in Professional Domains via Knowledge Augmented Generation" — arXiv:2409.13731

Your data has structure. It doesn't have meaning

Why enterprise AI has made a thirty-year-old question urgent — and what the move from Information Engineering, UML and relational modelling to graph, semantic and ontology-driven execution actually demands of practitioners.

The problem is not the model. It's the ground truth.

Shared meaning is not a new problem. Warehouse teams have been fighting "whose definition of revenue?" for thirty years. What changed is who is now reading the data.

A traditional report was read by an analyst who knew that "active customers" in the marketing dashboard was not the "active" the risk team used. She carried the caveat in her head, applied it, and moved on. The ambiguity was real but bounded — it lived in a few expert skulls and travelled no further than a meeting.

Put an agent in that seat. Ask it "how many active customers do we have in the North region, and what is their total exposure?" It will join the marketing table's active flag to the risk system's exposure column, sum the result, and return a confident number in fluent prose. It does not know the two systems mean different things by customer. It has no caveat. And its output looks exactly as authoritative as a correct answer would.

Three properties of AI-mediated work turn latent ambiguity into active cost:

  • Amplification. A human analyst makes one join an hour. An agent makes thousands, unsupervised. Every wrong assumption about what a field means is now executed at machine scale

  • Composition. Agents chain steps. A subtly wrong definition at step one propagates, and by step five the provenance of the error is gone. The answer is wrong and unexplainable at once

  • Opacity of confidence. The failure mode is not a crash or a NULL. It is a plausible, wrong, actionable answer - the expensive kind

The instinct, when the agent gets it wrong, is to reach for a bigger model or better prompts. But the error is not in the model's reasoning. It is in the ground truth the model reasoned over. If customer denotes four different things in four systems, no amount of model capability recovers the one you meant — because that thing was never written down anywhere a machine could read it.

Semantics is the missing input, not a missing capability.

"Semantic" means five different things

Here is where most programmes go wrong before they start. "Semantic layer" is the most overloaded phrase in enterprise data. Serious people use it to mean at least five distinct things that live in different places, are owned by different people, cost different amounts and solve different problems.

They are not rival definitions. They are a stack:

Read one word — customer — down that stack and the point becomes clear. Each reading is legitimate. Each is owned by a different team. Each fails in a different way.

Why this matters commercially: a retail bank funds a "Single View of Customer" programme. The deck talks about AI. The budget goes on a layer-3 analytics semantic layer, because that is what the vendor demonstrated. Eighteen months later the executive dashboard still shows three different customer counts — because the actual problem was layer 2. The same person exists as four unreconciled records in four systems, and no metrics layer reconciles identities it was never given. The programme built a precise calculator on top of an unresolved question.

Buying layer 3 to fix a layer-1 or layer-2 problem is the single most expensive mistake in this field.

Four demands, and the standards that answer them

1. Shared meaning — a common controlled vocabulary

Before anything else: do people agree what the words denote? "A customer is a natural or legal person who holds at least one live product with the bank." Plain language, agreed by humans, no code.

Most "data quality" disputes are vocabulary disputes wearing a technical costume. The formalisation is cheap and unglamorous: SKOS — the Simple Knowledge Organization System (W3C Recommendation, 2009) gives you concepts, preferred and alternative labels, broader/narrower hierarchies and mappings between schemes. ISO 25964 covers thesaurus construction and interoperability; ISO/IEC 11179 covers metadata registries and the discipline of naming and defining a data element properly.

This is the twenty per cent of the stack that needs no logic at all and returns most of the value. Skip it and everything above inherits the ambiguity.

2. Agreed derivation — the semantic layer for BI

Everyone agrees what a customer is, and two dashboards still disagree on the count, because the calculation lives in two places. The answer is a governed metric definition — active customer = at least one transaction in the last 90 days — written once, over the warehouse, and resolved identically by the dashboard, the notebook and the agent.

This is the sense most vendors mean today: the dbt Semantic Layer, Cube, AtScale, LookML. It is genuinely valuable, and it is also where the confusion concentrates, because it is the layer with the biggest marketing budget. A metrics layer computes flawlessly over whichever records you feed it. If those are four unreconciled customers, the metric is precise and wrong.

The real prize here is not the dashboard. It is that an agent can call a governed metric instead of misjoining raw tables.

3. Integration and interoperability — semantic expression

Point-to-point mapping does not survive past a handful of systems: n systems need n(n−1)/2 mappings, each with its own quiet assumptions. Meaning that travels needs to be expressed in a form that is not owned by any one system.

  • Semantic data is best expressed in graphs. A graph is a connected set of statements, conceptually nodes and edges. e.g. John -worksFor-> Enron -sells-> Oil

  • Graphs can be stored in simple fact statements (see RDF) or Labelled Property Graphs (see LPG)

  • RDF — the Resource Description Framework (W3C; RDF 1.1 in 2014, with RDF 1.2 reaching Candidate Recommendation in April 2026) reduces everything to the triple: subject, predicate, object. e.g. John worksFor Enron; Enron sells Oil. RDFS — RDF Schema adds classes, properties and subsumption (class subtyping)

  • IRIs — Internationalized Resource Identifiers replace local primary keys with globally resolvable names, so "the same customer" is a claim two systems can actually make

  • SPARQL — the SPARQL Protocol and RDF Query Language (1.1, 2013; 1.2 in draft) provides queries and updates across a graph of triples

  • RDF vs LPG — RDF data stores (otherwise known as Triple Stores) can accommodate data without schema (or with) and can expand easily to accommodate new data without structure change. They are best for integrating data from heterogeneous sources. Labelled Property Graphs are more efficient and best for transactional and operational data

  • R2RML — the Relational DB to RDF Mapping Language (W3C, 2012) and RML — the RDF Mapping Language lift the relational and file-based estate you already have into RDF form, declaratively, without rewriting the sources

  • JSON-LD — JavaScript Object Notation for Linked Data (1.1, W3C, 2020) carries the same meaning inside ordinary API payloads, which is how layer 4 gets done in practice

  • On the property-graph side, GQL — the Graph Query Language (ISO/IEC 39075:2024) is now an ISO standard in its own right, and SQL/PGQ (part of ISO/IEC 9075:2023) brings graph pattern matching into SQL. The RDF-versus-LPG argument is largely over: mature programmes run both and bridge between them

  • Reuse beats invention: There are many sources of well defined models addressing generic requirements as well as specific disciplines. Consult schema.org, FIBO — the Financial Industry Business Ontology, SNOMED CT — Systematized Nomenclature of Medicine, Clinical Terms, QUDT — Quantities, Units, Dimensions and Data Types, PROV-O — the PROV Ontology for provenance, DCAT — the Data Catalog Vocabulary for datasets

Using these technologies allows merging and reaching consensus on definition and holding that stable and shared in a standard format.

4. Reliability and agentic automation — formal ontology backed by logic

Here the stakes change. If an agent acts on your model, "roughly right" stops being a documentation problem and becomes an operational risk. That demands semantics formal enough for a machine to check.

  • OWL 2 — the Web Ontology Language (W3C Recommendation, 2012) sits on description logics, a decidable fragment of first-order logic. Its profiles — EL, QL, RL and the full DL — are deliberate trade-offs between what you can say and what a reasoner can compute in reasonable time. Choosing the profile is an architectural decision, not a detail

  • SHACL — the Shapes Constraint Language (W3C Recommendation, 2017; 1.2 drafts in progress) does the job OWL was never designed for: validating that actual data conforms to a stated shape. OWL says what things are; SHACL says what a valid record looks like. Conflating them is the most common beginner error in the field

  • Common Logic (ISO/IEC 24707:2018), and its CLIF — Common Logic Interchange Format syntax, expresses the constraints description logic deliberately cannot: n-ary relations, multi-property rules, conditions over time. Paired with a theorem prover and a model finder, it also lets you verify an ontology — state the intended models, prove the entailments you expect, and check that your axioms are independent and consistent. This is the discipline that separates an engineered ontology from a plausible-looking one

  • Top-level ontologies give your concepts somewhere to stand: BFO — the Basic Formal Ontology (ISO/IEC 21838-2:2021), DOLCE — the Descriptive Ontology for Linguistic and Cognitive Engineering, and gist (Semantic Arts small business high level ontology). OntoClean (Guarino & Welty) supplies the analytical tests — identity, rigidity, unity, dependence — that catch the taxonomy errors everyone else ships

Only at this rung does an agent get something it can be held to: classifications it can derive, constraints it can be blocked by, and an audit trail explaining why it concluded what it concluded.

The transition practitioners actually have to make

To be fair to the tools we already use, they carried us a very long way and none of this asks you to throw them out.

Codd's relational model (1970) gave data a rigorous mathematical foundation and the normalisation theory that removes update anomalies — arguably the most successful idea in applied computing. Chen's entity-relationship modelling (1976) gave us the picture: entities, attributes, cardinalities, drawn so a business person can read it and a database can be generated from it. UML class diagrams and the wider Information Engineering tradition added object structure and the discipline of conceptual → logical → physical design. The patterns literature — Hay, Simsion & Witt — is full of hard-won wisdom that remains correct: the Party pattern, role versus thing, resolving many-to-many relationships properly.

All of it comes with you. Graph modelling is not a rejection of data modelling. It is data modelling with the assumptions resolved.

The hard part is not the standards. It is the unlearning — and in my experience excellent data modellers find this transition harder than beginners do, precisely because the old instincts are so good. Thirty years of closed-world reflexes ("if there's no row, it didn't happen") do not switch off because you read the OWL primer. Training has to be designed around that, not around the specification.

Where to start

Most enterprises need a great deal of layers 1 to 3 and a careful, small amount of layer 5. Formal ontology is the most powerful and the most over-applied technique on the map. Reaching for OWL when a governed glossary and a resolved identity would have done is how semantic programmes acquire their reputation for cost and delay.

There is a newer failure mode too. Large language models will happily draft you an ontology in minutes. Without competency questions, OntoClean analysis, SHACL validation and reasoner checks, what you get is a vibe ontology: fluent, plausible, and unverifiable — the same failure as the confident wrong answer, one level further up the stack. Assisted, then verified. In that order.

What we're doing about it

These threads — shared vocabulary, the canonical model, the semantic layer, integration, formal ontology, and the bridge into agentic AI — are what we have spent this year building a training path around.

We're introducing it as two courses: a transition course for commercial analysts, BI practitioners and domain experts who need to become the bridge to semantic work; and an advanced engineering course for modellers, ontologists and AI implementors who have to build the solution and put it into production.

Ready to make this shift on your own data estate?

Course 1: Building Shared Meaning starts 19–23 October.

 

References and standards

Standards

  1. W3C, RDF 1.1 Concepts and Abstract Syntax, Recommendation, 2014 — w3.org/TR/rdf11-concepts; RDF 1.2 Concepts and Abstract Data Model, Candidate Recommendation, 2026 — w3.org/TR/rdf12-concepts

  2. W3C, SKOS Simple Knowledge Organization System Reference, Recommendation, 2009 — w3.org/TR/skos-reference

  3. W3C, OWL 2 Web Ontology Language Primer, 2nd edition, 2012 — w3.org/TR/owl2-primer

  4. W3C, Shapes Constraint Language (SHACL), Recommendation, 2017 — w3.org/TR/shacl; SHACL 1.2 Core, Working Draft — w3.org/TR/shacl12-core

  5. W3C, SPARQL 1.1 Query Language, Recommendation, 2013 — w3.org/TR/sparql11-query

  6. W3C, R2RML: RDB to RDF Mapping Language, Recommendation, 2012 — w3.org/TR/r2rml

  7. W3C, JSON-LD 1.1, Recommendation, 2020 — w3.org/TR/json-ld11

  8. ISO/IEC 39075:2024, Information technology — Database languages — GQL

  9. ISO/IEC 24707:2018, Information technology — Common Logic (CL)

  10. ISO/IEC 21838-2:2021, Top-level ontologies — Basic Formal Ontology (BFO)

  11. ISO 25964, Thesauri and interoperability with other vocabularies; ISO/IEC 11179, Metadata registries

Literature

  1. Codd, E.F. (1970), "A Relational Model of Data for Large Shared Data Banks", Communications of the ACM 13(6).

  2. Chen, P. (1976), "The Entity-Relationship Model — Toward a Unified View of Data", ACM TODS 1(1).

  3. Berners-Lee, T., Hendler, J. & Lassila, O. (2001), "The Semantic Web", Scientific American.

  4. Guarino, N. & Welty, C. (2009), "An Overview of OntoClean", in Handbook on Ontologies, Springer.

  5. Allemang, D., Hendler, J. & Gandon, F. (2020), Semantic Web for the Working Ontologist, 3rd edition, ACM Books.

  6. Hogan, A. et al. (2021), "Knowledge Graphs", ACM Computing Surveys 54(4).

  7. Noy, N. et al. (2019), "Industry-scale Knowledge Graphs: Lessons and Challenges", Communications of the ACM 62(8).

  8. Suárez-Figueroa, M.C. et al. (2012), Ontology Engineering in a Networked World (the NeOn methodology), Springer.

  9. Grüninger, M. & Fox, M. (1995), "Methodology for the Design and Evaluation of Ontologies", IJCAI Workshop on Basic Ontological Issues in Knowledge Sharing.

  10. Lewis, P. et al. (2020), "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", NeurIPS — arXiv:2005.11401

  11. Edge, D. et al. (2024), "From Local to Global: A Graph RAG Approach to Query-Focused Summarization" — arXiv:2404.16130

Graham McLeod · Inspired.org