E7 · Publication Volume 29
Safe Report Generation
claim–source mapping, limitations and versioned output
*claim–source mapping, limitations and versioned output*
Learning objectives and boundary
This lesson is general and institution-neutral. It uses no real company, individual, property, project or identifiable place. Generic roles describe responsibilities only, and every SYN-AI identifier denotes explicitly synthetic teaching evidence.
- Frame the decision governed by claim–source mapping, limitations and versioned output.
- Separate generated proposals from admissible evidence, deterministic results and accountable judgement.
- Define measurable failure, abstention, escalation and release conditions.
- Produce a versioned bilingual report package with complete claim–source verification from synthetic evidence and defend its controls.
Decision and professional boundary
Generate a report from approved atomic claims and evidence relations, not from an open request to “write the conclusion”. Structure precedes prose: define sections, permitted claim types, required source links, uncertainty language, limitations and sign-off. The generator may combine and paraphrase only within those constraints. Verification re-extracts atomic claims from the candidate and compares them with the approved claim set.
The report must distinguish observation, derived product, interpretation, hypothesis and recommendation. It states what is unknown, what evidence is excluded, which versions were used and which decisions remain outside scope. Safe generation improves readability without adding scientific authority.
Core concepts
| Concept | Operational meaning | |---|---| | Approved claim set | The versioned atomic propositions permitted to appear in the report. | | Claim–source map | A complete relation from each report claim to admissible supporting and contrary evidence. | | Limitation statement | A scoped description of missing evidence, method boundary or unresolved risk. | | Versioned report | An immutable candidate with source snapshot, generator configuration and review state. | | Semantic diff | A change report at claim, confidence, evidence and limitation level rather than text alone. |
Evidence model
The report package contains schema version, approved claim IDs, candidate prose, claim offsets, source locators, contradiction links, confidence records, required limitation IDs, generation configuration, verification results, semantic diff and approvals. A readable citation list is accompanied by machine-checkable mappings. Removed claims remain in history with removal reasons.
Controlled workflow
- Freeze the approved claim set, evidence graph and report schema.
- Assemble sections with observation, interpretation and hypothesis labels.
- Generate bounded prose with claim IDs retained internally.
- Re-extract candidate claims and compare with the approved set.
- Verify citations, contradictions, confidence language and required limitations.
- Review the semantic diff, sign the exact digest and publish immutably.
Every step emits a versioned artefact or an explicit failure. A later stage consumes only verified outputs from the preceding stage; conversational context is never an undocumented data channel.
Measures and acceptance criteria
Measure atomic support precision, citation completeness, contradiction disclosure, limitation coverage, semantic-diff accuracy and reviewer correction rate. Readability is assessed separately and cannot compensate for unsupported claims. A release gate requires every factual claim to be approved and sourced, every mandated limitation to appear and every high-consequence wording change to be reviewed.
$C_{citation}=\frac{N_{claims\ with\ sufficient\ citations}}{N_{factual\ claims}}$
| Gate | Required evidence | Release consequence | |---|---|---| | Identity | Resolvable claim, source, configuration and result IDs | Block when any identity is ambiguous | | Grounding | Every factual claim reaches sufficient source spans | Remove or withhold unsupported claims | | Independence | Lineage and partition checks show no circular or target information | Invalidate affected support and scores | | Review | Required roles disposition the exact candidate version | Keep the candidate non-published | | Reproducibility | Manifest, tools and checks reconstruct the evidence package | Return the package for correction |
Claim–source and system contract
The report contract defines section order, claim types, required evidence relations, citation style, uncertainty vocabulary, prohibited inferences, limitation IDs, language, accessibility requirements, verification thresholds and sign-off roles. The generator cannot create a new numeric value, coordinate, status or recommendation unless an approved claim explicitly contains it.
Security, privacy and access boundary
Render reports from a filtered evidence snapshot and scan outputs for restricted identifiers, hidden instructions and cross-access citations. Public versions are separate artefacts with explicit redaction records. Do not send protected evidence to an undeclared external service. Preserve the source package and report digest so later distribution can be verified.
Human review and escalation
Reviewers use a claim table and semantic diff, then inspect the rendered report for context, prominence and accessibility. A limitation hidden in an appendix may be technically present yet operationally ineffective. Sign-off binds to content digest, language, render and evidence snapshot; translation is reviewed as a new candidate because qualifiers can change.
Worked synthetic example
A synthetic report section has three approved claims: an observed alteration interval, an interpreted trend and a hypothesis about continuity. The generator labels each level and cites the interval record, interpretation snapshot and hypothesis matrix. Verification detects an added phrase “extends beyond the section”, which is absent from the approved set. The phrase is removed, and the required limitation “continuity beyond the section is untested” is placed beside the hypothesis.
Counterexample and failure analysis
A prompt includes raw documents and asks for a concise executive summary with citations. The output compresses qualifiers, merges current and obsolete sources and creates one unsupported recommendation. Citations at paragraph ends make it impossible to tell which source supports which claim. Safe reporting requires an approved claim layer and post-generation verification, not only a citation instruction.
Frequent failure modes
- Generating conclusions directly from raw context
- Attaching one citation to multiple ambiguous claims
- Dropping limitations during summarisation
- Treating translation as wording-only change
- Signing a report without its evidence snapshot
Practical exercise
- Create an approved claim set for a short synthetic report.
- Map every claim to supporting, contrary and missing evidence.
- Design a semantic diff that detects confidence and limitation changes.
- Write release checks for bilingual report versions.
Assessed artefact
Submit a versioned bilingual report package with complete claim–source verification with its source manifest, configuration versions, acceptance evidence, rejected cases, unresolved risks and a short explanation of why one plausible alternative control was not selected. The artefact is incomplete if its diagrams or prose cannot be reconciled with machine-checkable evidence.
Verification checkpoint
Trace one released synthetic claim backwards through review, confidence, verification, evidence graph, tool result or source span to an immutable input. Then trace one rejected claim to the earliest failed contract. Re-run the package from its manifest and compare content digests. The checkpoint passes only when a second reviewer can reconstruct the accepted and blocked paths, identify every assumption and reproduce the release decision without access to the original conversation. Record unresolved risk instead of completing a missing fact with generated text.
Sources
- FActScore atomic factual-precision evaluation, a primary method for decomposing long-form output into source-checkable atomic facts.
- FEVER fact-extraction and verification benchmark, a primary benchmark linking claims to supporting, refuting or insufficient evidence.
- PROV-O provenance ontology, a standard model for entities, activities, responsibility roles and derivation.
- Generative AI Risk Management Profile, a cross-sector profile covering risks and controls specific to generative systems.