E7 · Publication Volume 29
AI, LLMs, Evidence Graphs and Verifiable Geoscience Reasoning
A general, institution-neutral tutorial for grounded, tool-mediated and human-reviewed AI reasoning tied to explicit geoscience evidence, with no affiliation to any company, individual, model vendor or technology platform.
Purpose and institutional neutrality
A general, institution-neutral tutorial for designing evidence-grounded, tool-mediated and human-reviewed AI workflows for geoscience.
This is a general, institution-neutral tutorial. It has no relationship to any company or individual. Every worked example uses the explicitly synthetic SYN-AI evidence package, generic responsibility roles and fictitious identifiers; none describes a real property, project, organisation or identifiable place. External people and institutions are named only in Sources so readers can trace the public standards and primary research used to prepare the material.
The delivery website is only a host. It is not the publisher, scientific authority or curriculum subject. Hosting does not create authorship, sponsorship, endorsement, employment or affiliation. Claims in this tutorial are governed by declared evidence, reproducible transformations, evaluation and accountable human review.
Audience and prerequisites
This volume is for geoscientists, data professionals, software engineers, reviewers and students who need to evaluate or design AI-assisted evidence workflows. It assumes the data semantics and provenance foundations of E2, the visual review concepts of E5, the architecture and testing foundations of E6, and relevant domain knowledge from the geological volumes. No particular model, vendor, cloud, database or programming language is required.
Readers should be able to distinguish observation from interpretation, inspect a data manifest, reason about uncertainty and recognise that a generated sentence is not evidence. Programming helps with exercises, but the essential work is specifying boundaries, evidence, tests and review.
Learning outcomes
- Select appropriate bounded AI tasks and reject uses that cross professional or evidential boundaries.
- Build resolvable grounding and retrieval packages for text, tables, figures and structured data.
- Design typed tool execution, evidence graphs and competing-hypothesis reasoning.
- Measure claim support, calibration, abstention, leakage and holdout generalisation.
- Construct meaningful human review, safe report generation and continuous evaluation.
- Deliver a complete, versioned and auditable claim–source reasoning package.
Verification architecture
The tutorial uses one recurring chain: request → task contract → grounded evidence → retrieval and admission → typed tools → evidence graph → competing hypotheses → candidate claims → claim verification → calibrated confidence → human disposition → versioned report → monitoring. Every arrow is a controlled transformation with recorded inputs, outputs and failure states.
Language generation never sits at the source of scientific truth. It can propose, organise and explain within an evidence boundary. Deterministic tools calculate; source objects provide observations; reviewers judge scope and consequences. Abstention and escalation are valid outputs whenever required evidence or authority is missing.
Synthetic case package
SYN-AI is a deliberately fictitious evidence package containing synthetic interval descriptions, maps, raster summaries, document pages, claims, tool results, hypotheses, review events and holdout outcomes. Values are selected to expose ambiguity, duplicated lineage, inconsistent versions, spatial dependence and missing evidence. They are not estimates for a real deposit or operational system.
Each exercise begins from an immutable manifest. Learners may change retrieval, prompts, policies, tools or thresholds only by issuing a new version. Hidden evaluation objects are kept outside development paths so the same package can demonstrate leakage and blind validation without referring to a real organisation.
Curriculum map
- E7-01 — Appropriate and Inappropriate Uses of AI. extraction, classification, search and reasoning limitations
- E7-02 — Document and Data Grounding. chunking, tables, figures, metadata and citation
- E7-03 — Retrieval and Evidence Selection. recall, precision, source quality and temporal relevance
- E7-04 — Tool-Mediated Geoscience Agents. runtime tools, network boundaries and deterministic computation
- E7-05 — Evidence-Graph Models. observation, product, interpretation, hypothesis and dependency
- E7-06 — Competing Hypotheses and Structured Reasoning. predictions, support, contradiction and missing evidence
- E7-07 — Confidence and Calibration. claim-level confidence, source confidence and uncertainty language
- E7-08 — Hallucination, Leakage and Circular Evidence. self-citation, derived-data reuse and target leakage
- E7-09 — Human Review and Responsibility. review roles, approval, escalation and audit trail
- E7-10 — Blind and Holdout Validation. frozen predictions, independent evidence and scoring
- E7-11 — Safe Report Generation. claim–source mapping, limitations and versioned output
- E7-12 — Monitoring and Continuous Evaluation. drift, benchmark sets, failure taxonomy and feedback
Assessment and completion outcome
Every lesson produces an evidence-bearing artefact: a suitability register, grounding package, retrieval benchmark, controlled runtime, evidence graph, hypothesis matrix, calibration report, integrity audit, review protocol, holdout report, versioned report or monitoring plan. Assessment checks reproducibility, failure handling and professional boundaries, not writing style alone.
The completion outcome is a small but complete geoscience reasoning workflow whose claims can be traced to source spans and tool results, whose alternatives and unknowns remain visible, whose confidence is calibrated on independent evidence and whose final report has an accountable review history.
Version and source policy
AI behaviour, evaluation practice and standards evolve. Every implementation record must identify the exact source edition, corpus snapshot, tool contract, model configuration, policy and benchmark version used. A newer model or specification does not silently update an approved result; it creates a candidate change that must pass the same evidence and review gates.
Public standards and primary research are used as technical evidence, not as ownership or endorsement. Their names and links appear only in Sources. The tutorial remains portable across implementations because contracts and acceptance evidence are defined independently of products.
Sources
- AI Risk Management Framework 1.0, a lifecycle framework for governing, mapping, measuring and managing context-dependent AI risk.
- Generative AI Risk Management Profile, a cross-sector profile covering risks and controls specific to generative systems.
- ISO/IEC 23894:2023 AI risk-management guidance, guidance for integrating AI-specific risk management into activities and decisions.
- PROV-O provenance ontology, a standard model for entities, activities, responsibility roles and derivation.