E7 · Publication Volume 29

Evidence-Graph Models

observation, product, interpretation, hypothesis and dependency

*observation, product, interpretation, hypothesis and dependency*

Typed evidence graph separating observations, products, interpretations, claims and hypotheses
Typed evidence graph separating observations, products, interpretations, claims and hypotheses

Learning objectives and boundary

This lesson is general and institution-neutral. It uses no real company, individual, property, project or identifiable place. Generic roles describe responsibilities only, and every SYN-AI identifier denotes explicitly synthetic teaching evidence.

  • Frame the decision governed by observation, product, interpretation, hypothesis and dependency.
  • Separate generated proposals from admissible evidence, deterministic results and accountable judgement.
  • Define measurable failure, abstention, escalation and release conditions.
  • Produce a validated, versioned geoscience evidence-graph model and snapshot from synthetic evidence and defend its controls.

Decision and professional boundary

An evidence graph records what exists, how it was produced and what each claim depends on. It is not a decorative network and it does not convert connections into truth. Separate source objects, observations, derived products, interpretations, atomic claims, hypotheses, tests and review decisions. Use typed, directed edges such as derived from, observes, supports, contradicts, qualifies, depends on and supersedes. Every edge has scope, version and responsibility.

Truth status is claim specific. One observation may support a lithological description while saying nothing about continuity. A derived product may be reproducible yet unsuitable for a particular scale. The graph must preserve negative and missing evidence, alternative interpretations and invalidated dependencies instead of showing only the preferred story.

Core concepts

| Concept | Operational meaning | |---|---| | Evidence node | An identifiable source, observation or derived product that can be inspected. | | Claim node | One bounded proposition with subject, predicate, object, scope and status. | | Support edge | A reviewed relation stating how evidence bears on one claim under declared conditions. | | Dependency edge | A relation showing that validity relies on another object, activity or assumption. | | Validity interval | The time or version range during which a node or relation is applicable. |

Evidence model

Use immutable evidence nodes and versioned assertion nodes. A relation record contains source node, target node, relation type, applicability, rationale, activity, reviewer state and confidence basis. Distinguish an observation made in the world from a digital representation of that observation and from an interpretation derived from it. This prevents a copied map feature from masquerading as independent field evidence.

Controlled workflow

  1. Define node and edge types before loading instances.
  2. Assign immutable identities and versions to evidence objects.
  3. Decompose interpretations into atomic scoped claims.
  4. Link support, contradiction, qualification and dependency with rationale.
  5. Validate graph shapes, prohibited cycles and required provenance paths.
  6. Materialise review views without deleting alternatives or history.

Every step emits a versioned artefact or an explicit failure. A later stage consumes only verified outputs from the preceding stage; conversational context is never an undocumented data channel.

Measures and acceptance criteria

Measure provenance-path completeness, orphan rate, unresolved-identifier rate, prohibited-cycle count, support-rationale completeness and alternative-retention coverage. Graph density is not a quality measure. A release view passes only when every visible claim reaches an inspectable evidence node or an explicit unknown state, and every derived product reaches its source and transformation activity.

$C_{prov}=\frac{N_{claims\ with\ complete\ provenance\ paths}}{N_{released\ claims}}$

| Gate | Required evidence | Release consequence | |---|---|---| | Identity | Resolvable claim, source, configuration and result IDs | Block when any identity is ambiguous | | Grounding | Every factual claim reaches sufficient source spans | Remove or withhold unsupported claims | | Independence | Lineage and partition checks show no circular or target information | Invalidate affected support and scores | | Review | Required roles disposition the exact candidate version | Keep the candidate non-published | | Reproducibility | Manifest, tools and checks reconstruct the evidence package | Return the package for correction |

Claim–source and system contract

The graph contract defines allowed node classes, edge direction, cardinality, required identifiers, provenance depth, validity semantics, review states and cycle rules. Support and contradiction edges require a rationale and cannot originate from generated prose alone. A summary view is a query over the graph, never a replacement for the graph.

Security, privacy and access boundary

Authorisation is evaluated on traversals, not just nodes. A public claim must not reveal a restricted predecessor through labels, counts or inferred paths. Generate filtered subgraphs with explicit redaction nodes so absence caused by access control is not confused with absence of evidence. Protect relation history because deleted or rewritten edges can change conclusions without changing source files.

Human review and escalation

Review samples of complete paths, negative evidence and cross-version differences. Domain reviewers judge whether a relation is scientifically meaningful; data reviewers check identity, provenance and constraints. Approval names a graph snapshot and query, because the same graph can support different views under different filters.

Worked synthetic example

SYN-AI records two synthetic outcrop observations, one raster product and two interpretations of a contact. The raster derives from observations and interpolation parameters; interpretation A follows a topographic break, while interpretation B follows sparse structural measurements. A new observation contradicts one local segment of A but does not invalidate the whole interpretation. The graph records scoped contradiction, keeps B, and marks continuity beyond the new observation as unknown.

Counterexample and failure analysis

A graph that links every document mentioning “fault” to one fault concept creates semantic proximity, not evidence. It omits whether the document reports an observation, repeats an interpretation or rejects the term. Counting links then rewards duplication and can amplify circular evidence. Edge types and rationales must express epistemic meaning.

Frequent failure modes

  • Treating any link as support
  • Collapsing observations and interpretations into one node type
  • Deleting superseded or contradictory relations
  • Counting duplicated lineage as corroboration
  • Publishing a filtered graph without redaction semantics

Practical exercise

  1. Model one observation, product, interpretation, claim and hypothesis as distinct nodes.
  2. Write shape constraints for a support edge.
  3. Find and explain a prohibited evidence cycle.
  4. Create a version diff that preserves an invalidated relation.

Assessed artefact

Submit a validated, versioned geoscience evidence-graph model and snapshot with its source manifest, configuration versions, acceptance evidence, rejected cases, unresolved risks and a short explanation of why one plausible alternative control was not selected. The artefact is incomplete if its diagrams or prose cannot be reconciled with machine-checkable evidence.

Verification checkpoint

Trace one released synthetic claim backwards through review, confidence, verification, evidence graph, tool result or source span to an immutable input. Then trace one rejected claim to the earliest failed contract. Re-run the package from its manifest and compare content digests. The checkpoint passes only when a second reviewer can reconstruct the accepted and blocked paths, identify every assumption and reproduce the release decision without access to the original conversation. Record unresolved risk instead of completing a missing fact with generated text.

Sources