Knowledge engineering by domain
A bounded domain is the only scope that finishes. These pages give one way in per domain: the published backbone to reuse, the questions to agree first, the smallest artifact worth building, and the failure the domain is prone to.
Why a domain, and not the whole organization
Knowledge engineering fails more often from scope than from tooling. A programme that sets out to model the enterprise before proving value on one use case tends to stall at the specification stage, having produced a document nothing runs on. A bounded domain is the only scope that finishes.
There is a structural reason beyond project management. The April 2026 preprint cited below organizes its domain layer hierarchically — an agent working in fintech.payments inherits the concepts of fintech and adds the payment-specific ones — and uses that hierarchy to decide which tools an agent may even see. Model a bounded domain and the inheritance works. Model everything and there is nothing to inherit from.
This is also where a staged approach pays. The ontology pipeline set out by Jessica Talisman in the talk cited below is explicit that each stage is usable on its own: an organization with a controlled vocabulary and private metadata schemas may have all the context it needs for now. A domain playbook is a way into that ladder, not a commitment to climb it all.
The method, the same on every page
Each domain page answers six questions in the same order, so the pages can be compared rather than merely collected.
- What to reuse. The published standard, vocabulary or model that covers most of the domain, named with its number and year where it has one.
- What to scope to for a first attempt — one reporting line, one clinical pathway, one repository — rather than the whole domain.
- Which questions to agree first: ten to twenty concrete questions the knowledge must be able to answer, settled with domain experts before any modelling begins.
- The smallest artifact worth building, which is almost always a vocabulary of terms with definitions and stewards, plus a small set of hard constraints.
- How to validate it, against a system of record or against the domain experts themselves rather than against a demonstration.
- The failure the domain is prone to — the specific way an agent goes wrong here without raising an error.
The ten
| Domain | Backbone to reuse | Kind |
|---|---|---|
| Library and information science | Z39.19, ISO 25964, SKOS, IFLA LRM, BIBFRAME, VIAF, OAI-PMH | Horizontal, origins — written |
| Finance and insurance | FIBO, ISO 20022, XBRL, Basel III and IV terminology | Vertical — in preparation |
| Healthcare and life sciences | SNOMED CT, LOINC, HL7 FHIR, ICD, DICOM, OMOP CDM | Vertical — in preparation |
| Security | MITRE ATT&CK, CVE and CVSS, CWE, STIX/TAXII, NIST CSF | Vertical — in preparation |
| Aerospace and aviation | S1000D, AIXM, ATA iSpec 2200 | Vertical, effectivity-hard — in preparation |
| Software engineering | SPDX, CycloneDX, OSV, SLSA, CWE, OpenAPI and AsyncAPI | Vertical, freshness-hard — in preparation |
| HR and workforce | ESCO, O*NET, ISCO-08, HR Open Standards | Vertical — in preparation |
| Industrial: supply chain, manufacturing, energy | GS1 and EPCIS, IEC 62264, Asset Administration Shell, AutomationML, IEC CIM | Vertical, clustered — in preparation |
| Legal and regulatory | Akoma Ntoso, LegalRuleML, jurisdiction models | Vertical, versioning-hard — in preparation |
| Enterprise business, general | schema.org, gist, BORO, REA, upper ontologies | Horizontal — written |
Two of the ten are written. The other eight are named with their backbone so the shape of the set is visible, and no link is published until the page behind it exists.
Which one to start with
Not the domain the model already handles well. The preprint cited below reports that the gain from ontological grounding is largest where the model's pretrained knowledge is weakest, which it calls the inverse parametric knowledge effect. That inverts the usual instinct: a general-purpose domain is the worst first choice, because the model sounds competent there and the errors are quiet.
Two practical tests for a first domain. Is there a published backbone to align to? That decides the cost. And is there a system of record to validate against? That decides whether you can tell whether you succeeded.
What these pages do not do
- They do not rank vendors, and they do not score an organization's maturity. A score that nobody acts on is a report, not a result.
- They do not link a standard they could not verify from an automated client. The W3C, the Library of Congress, IFLA and ISNI answer such requests with a Cloudflare challenge; those are cited by name, number and year instead.
- They do not present the preprint's numbers as settled. It is one preprint, self-reported, and it is labelled as such every time it is cited.
FAQ
Ten are listed. Two are written in full — library and information science, and enterprise business general — because between them they hold the method the other eight pages point back to. The remaining eight are marked in preparation and are not linked, so nothing here leads to a missing page.
Because for most domains the reusable work already exists: FIBO for financial instruments, SNOMED CT and LOINC for clinical terms, MITRE ATT&CK for adversary behavior, GS1 for supply-chain items. The ontology task is alignment and extension, not construction. Build from scratch only for the genuinely organization-specific slice.
No. The clearest quantitative result is a preprint from April 2026 reporting 600 controlled runs across four sectors, and it is self-reported by the platform's own authors. It is cited here as a preprint with a date, not as an established result. The reuse argument above does not depend on it.
Because they cannot be verified from an automated client. The W3C, the Library of Congress, IFLA and ISNI answer such requests with a Cloudflare challenge rather than the page, and ISO is not linkable at all. Those are cited by name, number and year, following the convention already used for ISO on this site.
Sources
Standards published by the W3C, the Library of Congress, IFLA, NLM and ISNI are cited by name and date rather than linked: those hosts answer automated clients with a challenge page or a redirect stub rather than the document, so a link could not be verified. ISO standards are cited by number for the same reason. Every link above was fetched and its body inspected before publication.
- ANSI/NISO Z39.19-2005 (R2010), Guidelines for the Construction, Format, and Management of Monolingual Controlled Vocabularies
- W3C, SKOS Simple Knowledge Organization System Reference, W3C Recommendation, 18 August 2009 (cited by name and date: w3.org answers automated clients with a Cloudflare challenge)
- ISO 25964-1:2011, Information and documentation — Thesauri and interoperability with other vocabularies, Part 1 (cited by number; ISO pages are not linkable)
- IFLA, Library Reference Model (IFLA LRM), 2017 (cited by name and date; ifla.org is not reachable from an automated client)
- Library of Congress, BIBFRAME 2.0 (cited by name; loc.gov is not reachable from an automated client)
- NLM, Medical Subject Headings (MeSH) (cited by name; the MeSH home path answers automated clients with a redirect stub)
- Getty Research Institute, Art & Architecture Thesaurus
- Open Archives Initiative, Protocol for Metadata Harvesting, version 2.0
- DCMI Metadata Terms
- CIDOC CRM, the cultural-heritage reference model, published as ISO 21127
- ORCID, researcher identifiers
- DOI Foundation, the DOI system
- Semantic Arts, gist — an upper ontology for business
- schema.org, a shared vocabulary for describing things
- Wilkinson et al., The FAIR Guiding Principles for scientific data management and stewardship, Scientific Data 3, 2016
- Tuan, Ontology-Constrained Neural Reasoning in Enterprise Agentic Systems, arXiv 2604.00555, April 2026 — a preprint, self-reported by the platform's own authors
- Jessica Talisman, Knowledge Infrastructures and the Ontology Pipeline for AI Systems (O'Reilly Data Superstream, video), transcript of 0:00-18:56
Page updated .