E2 · Publication Volume 24

Controlled Vocabularies and Concepts

codes, synonyms, SKOS-compatible concepts and governance

Learning objectives

  • Explain why codes, synonyms, SKOS-compatible concepts and governance require explicit semantic modelling.
  • Design identities, relations and constraints that preserve controlled vocabularies and concepts across exchange.
  • Separate hard release gates from diagnostic metrics and interpretation choices.
  • Produce a versioned bilingual concept scheme with governed legacy mappings from synthetic evidence.

The lesson is complete only when the learner can defend both the model and the release decision. A neat schema without evidence, tests or declared limitations is an unverified design. The assessed artefact must make assumptions visible and distinguish source assertions from derived conclusions.

Decision context

Use controlled concepts when values need consistent comparison, validation or exchange. Do not turn free descriptive notes into arbitrary codes merely for storage convenience. The scheme must define scope, preferred labels, definitions, status, relationships and change policy.

Start with a decision record: name the intended use, the evidence required, the consequence of error, the accepted uncertainty and the role authorised to accept residual risk. Then ask whether the proposed model can answer the decision question without relying on filename conventions, row order, undocumented defaults or someone’s memory. This prevents technology selection from concealing a missing semantic requirement.

The same record may be fit for one use and unfit for another. A rapid exploratory view can tolerate conditions that a released exchange package cannot. Fitness is therefore stated against a use, contract version and quality gate rather than attached permanently to the data.

Core concept

A code is a local token; a concept is the meaning that token denotes. Separating them allows many labels, languages and legacy codes to refer to one durable concept. Broader, narrower and related links make the concept scheme navigable without pretending that every relationship is exact equivalence.

The working scope is codes, synonyms, SKOS-compatible concepts and governance. For each item in that scope, distinguish the thing itself, the label used by a source, the claim made about it and the record that carries the claim. Identity is not a display name; a value is not its unit; an observation is not a model; current is not the same as valid. These distinctions create explicit places for correction, uncertainty and competing interpretations.

A good semantic design can be explained as a set of sentences before it is encoded. Each sentence identifies a subject, a property or relationship, an object or result, and the context under which the claim holds. Physical tables and files are then projections of those sentences, not their source of meaning.

Semantic model

Each concept receives a persistent identifier independent of its preferred label. Attach one preferred label per language, any number of alternative labels, a definition, notes, status and concept-scheme identifier. Use typed semantic relations. Mapping to another scheme records exact, close, broader, narrower or related correspondence and the mapping authority.

Test every proposed record against seven questions: What has identity? What type is it? Which property or relationship is asserted? Which spatial and temporal context applies? Which state or qualifier modifies the assertion? Which evidence supports it? Which version and activity produced the stored representation? Missing answers become explicit contract gaps.

Normalisation is used to separate independent facts, not to maximise the number of tables. A compact nested object can be semantically sound if the same identities, constraints and provenance remain explicit. Conversely, a highly normalised database can still be ambiguous when relationships and units exist only in documentation or application code.

Constraints and invariants

| Invariant | Executable or review test | | --- | --- | | Concept identifiers persist | Renaming a preferred label does not create a new concept. | | Labels are language-tagged | Allow one preferred label per concept and language. | | Mappings have relation types | Do not collapse close, broader or narrower mappings into equality. | | Deprecated concepts remain resolvable | Supply status, replacement guidance and effective dates. |

An invariant is a condition that must remain true across storage, export, correction and reprocessing. Implement it as close to the authoritative boundary as practical and repeat the check at exchange boundaries. Record rule identifier, version, severity, evaluated scope, observed value and outcome so a failure can be reproduced.

Hard gates protect identity, semantic validity, required provenance and authorised use. Diagnostic checks reveal unusual values or patterns but require interpretation. Never convert a diagnostic threshold into deletion or correction without a reviewed rule and preserved source evidence.

Quantitative reasoning

Measure usage coverage C_v = n_m / n_t, where n_m is the number of values mapped to active concepts and n_t is the number of controlled values. Also report ambiguous mappings, deprecated usage and concepts lacking definitions. High coverage with many forced inexact mappings is not semantic quality.

Every reported ratio states its numerator, denominator, exclusions and evaluation time. Stratify results by source, entity type, contract version or processing run where aggregation could hide a local failure. Counts accompany percentages so a seemingly large change based on a tiny denominator remains visible.

Precision is part of meaning. Do not add decimal places merely because a storage type permits them, and do not round identity, interval or coordinate fields without a declared tolerance and test. Quantitative summaries support a release decision; they do not replace semantic review.

Evidence and uncertainty

Every mapping is a claim supported by definitions, examples and domain review. Preserve mapping provenance and confidence. Lexical similarity is useful for proposing candidates but cannot determine conceptual equivalence; homonyms and differences in scope make label-only matching unsafe.

Build an evidence packet containing preserved source reference, acquisition or assertion context, applicable method, validation results, reviewer decision and links to derivatives. Classify uncertainty as observational, semantic, structural, parametric or policy-related where that distinction changes treatment. “Unknown” is a valid state when the evidence cannot justify a stronger claim.

Contradictory evidence remains available. The model may select one current assertion, but the reason, competing assertion and effective time are retained. This makes later reinterpretation possible without pretending the earlier evidence never existed.

Interfaces and storage

Exchange concept identifiers in data records and distribute the scheme as a separately versioned resource. Include labels for display, but do not make consumers join by label. A contract states the accepted scheme and version range, plus behaviour for unknown or deprecated concepts.

Design an interface from the logical contract outward. Specify identifiers, types, cardinalities, units, value states, coordinate and time references, version negotiation, validation behaviour and structured errors before choosing a serialisation. The physical representation then declares its mapping to those logical elements.

Storage optimisation may partition, compress, index or cache data, but it must not change identity or silently remove context. A derived representation points to immutable inputs and a processing manifest. A cache carries freshness and contract-version information and is never treated as the only evidence copy.

Governance and access

A change proposal identifies the problem, affected concepts, compatibility impact and migration guidance. Review separates editorial label changes from semantic changes. Releases are immutable, machine-readable and accompanied by differences so producers and consumers can plan adoption independently.

Governance is expressed through named roles, review states and versioned decisions, not through references to a particular organisation. Define who may propose, validate, approve, supersede and withdraw each governed resource. The audit trail records the role and event while avoiding unnecessary personal data.

Apply least-necessary access to source evidence and derivatives. Access controls must not erase identifiers, lineage or quality metadata needed to understand an authorised release. When policy is unresolved, quarantine the output with a precise reason and escalation route.

Integration checkpoint

Concept identity, labels, hierarchy and mapping relations
Concept identity, labels, hierarchy and mapping relations

The diagram summarises the control flow for this lesson. Read it from source evidence through semantic structure and validation to a decision-ready artefact. Each arrow should correspond to a declared relationship or transformation; each boundary should have a contract; each released node should have an identity, version and provenance pointer.

Integrate the lesson by adding a versioned bilingual concept scheme with governed legacy mappings to the evolving synthetic data package. Verify that earlier artefacts still resolve and that the new model does not overwrite observations, identifiers, values or versions introduced in previous lessons. Record every changed assumption.

Synthetic worked example

A synthetic lithology list contains “mafic volcanic”, “basic volcanic” and a local code. Definitions show that two labels are synonyms in one scheme, while a third term has broader scope. The learner assigns one exact alias mapping, one broader mapping and leaves an unsupported candidate unresolved rather than forcing equivalence.

Work the example in four passes:

  1. Preserve the received records and write the intended decision without correcting anything.
  2. Identify entities, claims, context, uncertainties and policy constraints; mark every unresolved item.
  3. Apply the versioned rules, create derivatives and record the exact transformation plus validation evidence.
  4. Issue an accept, reject or quarantine decision and show how an independent reviewer can reproduce it.

Because the example is entirely synthetic, its values demonstrate method only. The important result is the chain from received evidence to justified decision. If a required fact is absent, the worked solution records the gap rather than manufacturing a plausible value.

Practice task

Create a bilingual concept scheme for twelve synthetic terms. Give every concept a stable identifier, definition, language-tagged labels, hierarchy and status. Map five legacy codes with typed relations, then validate preferred-label uniqueness, relationship targets and deprecated replacements.

Use the following acceptance criteria:

  • All required identifiers and references resolve to declared types.
  • Every transformation preserves the received evidence and records its derivation.
  • Invalid, unknown and inapplicable states remain distinct and machine-testable.
  • The output identifies the contract, vocabulary and processing versions used.
  • A second reader can reproduce the validation result without private knowledge.

Submit the source snapshot, authored contract or model, validation output, derivative, manifest and a short decision record. A screenshot alone is insufficient because it cannot demonstrate the exact input, version or rule execution.

Common failure modes

  • Codes and human-readable labels are treated as the same thing.
  • Every cross-scheme mapping is declared exact.
  • A renamed label creates a new concept identity.
  • Deprecated concepts disappear and make historical data unreadable.

These failures share a pattern: convenient representation is mistaken for verified meaning. Diagnose the earliest boundary at which an assumption became implicit. Correct by restoring source evidence, making the assumption a versioned field or rule, rerunning dependent transformations and superseding—not overwriting—the affected release.

Do not repair a failure by adding an undocumented default. A blocked result with a specific missing dependency is safer and more reusable than a complete-looking result whose meaning cannot be reconstructed.

Review questions

  1. How does a concept differ from a code?
  2. Why are language tags part of label semantics?
  3. When is a broader mapping safer than an exact mapping?
  4. What must accompany concept deprecation?

For each answer, identify the governing invariant, the evidence needed to evaluate it and the appropriate release behaviour when the invariant fails. A strong answer distinguishes scientific uncertainty from missing semantics and distinguishes a recoverable warning from a hard contract violation.

Sources and further reading

  • W3C SKOS Reference, a model for concepts, labels, broader relations and mapping.
  • OGC GeoSciML 4.1, an exchange model for geologic features, observations and vocabularies.
  • DCMI Metadata Terms, general resource-description properties and controlled ranges.
  • W3C PROV-O, a formal vocabulary for entities, activities, agents and derivation.