E7 · Publication Volume 29
Appropriate and Inappropriate Uses of AI
extraction, classification, search and reasoning limitations
*extraction, classification, search and reasoning limitations*
Learning objectives and boundary
This lesson is general and institution-neutral. It uses no real company, individual, property, project or identifiable place. Generic roles describe responsibilities only, and every SYN-AI identifier denotes explicitly synthetic teaching evidence.
- Frame the decision governed by extraction, classification, search and reasoning limitations.
- Separate generated proposals from admissible evidence, deterministic results and accountable judgement.
- Define measurable failure, abstention, escalation and release conditions.
- Produce an AI task-suitability and professional-boundary register from synthetic evidence and defend its controls.
Decision and professional boundary
Use AI only after defining the decision, the admissible evidence and the consequence of error. Extraction from a declared table, classification against a reviewed vocabulary and search over a bounded corpus can be suitable when outputs are checked against source records. Open-ended geological judgement, safety-critical approval, public-resource statements or claims whose evidence cannot be inspected are not delegated to a language model. The system may assemble evidence and propose alternatives; accountable domain judgement remains outside the model boundary.
Suitability is a property of the whole task configuration, not of a model label. The same capability may be low consequence when drafting search terms and high consequence when changing a geological domain boundary. Assess reversibility, detectability of error, evidence completeness, required competence and the availability of independent review. If a deterministic rule, database constraint or ordinary search solves the task, that simpler control is the baseline.
Core concepts
| Concept | Operational meaning | |---|---| | Decision support | Evidence is organised for a human decision; the system does not acquire professional authority. | | Bounded task | Inputs, outputs, vocabulary, corpus, tools and failure consequences are explicitly limited. | | Deterministic baseline | A simpler reproducible method establishes whether AI adds measurable value. | | Abstention | The safe output is no conclusion when evidence or control requirements are not met. | | Consequence class | Review and release controls scale with plausible harm, irreversibility and exposure. |
Evidence model
Represent the task as a chain from request to evidence set, controlled transformation, candidate output, verification result and human disposition. Every transition records input identity, version, rule or model configuration, tool calls and rejection reasons. A useful result is not merely fluent; it preserves enough state for a reviewer to reconstruct why the task was accepted, how the answer was bounded and which conditions would have forced abstention.
Controlled workflow
- Write the supported decision and explicit non-decisions.
- List admissible evidence, unavailable evidence and required domain competence.
- Implement a deterministic or manual baseline before adding probabilistic behaviour.
- Classify consequence, reversibility, exposure and error detectability.
- Define abstention, escalation and independent-review triggers.
- Compare evidence quality, error modes and cost against the baseline.
Every step emits a versioned artefact or an explicit failure. A later stage consumes only verified outputs from the preceding stage; conversational context is never an undocumented data channel.
Measures and acceptance criteria
Measure value against the baseline with task-specific correctness, evidence coverage, abstention quality and review burden. Report separate rates for accepted, rejected and escalated cases. A high average score does not justify use if severe errors cluster in a consequential subgroup. Acceptance requires zero unsupported release claims in the synthetic high-consequence set and a recorded disposition for every abstention.
$R_{task}=P(error)\times C(error)\times E(exposure)$
| Gate | Required evidence | Release consequence | |---|---|---| | Identity | Resolvable claim, source, configuration and result IDs | Block when any identity is ambiguous | | Grounding | Every factual claim reaches sufficient source spans | Remove or withhold unsupported claims | | Independence | Lineage and partition checks show no circular or target information | Invalidate affected support and scores | | Review | Required roles disposition the exact candidate version | Keep the candidate non-published | | Reproducibility | Manifest, tools and checks reconstruct the evidence package | Return the package for correction |
Claim–source and system contract
The task contract names the decision owner role, permitted operations, corpus boundary, output schema, required citations, maximum consequence class, refusal conditions and review state. It explicitly says that generated text is a candidate artefact. A release gate checks the contract independently of the generating component.
Security, privacy and access boundary
Start with read-only access to the minimum evidence package. Do not expose credentials, unrestricted networks, unpublished personal data or write-capable operational tools to an exploratory workflow. Inputs from documents are untrusted content, not instructions. Consequential actions require a separately authorised controller and explicit human approval.
Human review and escalation
A reviewer first checks task suitability, then evidence and output. If the task crosses a professional boundary, review stops before wording quality is considered. Approval records the exact candidate version and does not transfer responsibility to the system.
Worked synthetic example
SYN-AI contains 240 synthetic interval descriptions and a reviewed six-class alteration vocabulary. The baseline uses exact terms and deterministic precedence rules. A candidate classifier may suggest a class only when it quotes the decisive phrase, reports the vocabulary version and passes a contradiction check. One interval says “weak pale mica; uncertain origin” and lacks the required observation confidence. The model proposes a class, but the contract withholds it and routes the interval to review. The useful behaviour is the controlled abstention, not the plausible label.
Counterexample and failure analysis
A polished narrative that ranks drilling targets directly from mixed notes is inappropriate even if each sentence sounds geological. The request hides competing objectives, missing costs, inconsistent coordinate references and professional approval. Adding a disclaimer after generation does not repair the unbounded decision. The task must be decomposed into evidence extraction, explicit hypotheses, deterministic calculations and accountable review.
Frequent failure modes
- Using fluency as evidence of correctness
- Letting an experimental workflow write operational records
- Evaluating only average accuracy
- Calling a disclaimer a control
- Treating model refusal as a complete risk strategy
Practical exercise
- Rewrite an unbounded “interpret this prospect” request as three bounded support tasks.
- Construct a suitability matrix for extraction, classification, search and professional judgement.
- Define a deterministic baseline and one measurable reason to retain an AI component.
- Design an abstention record that a reviewer can audit without the original conversation.
Assessed artefact
Submit an AI task-suitability and professional-boundary register with its source manifest, configuration versions, acceptance evidence, rejected cases, unresolved risks and a short explanation of why one plausible alternative control was not selected. The artefact is incomplete if its diagrams or prose cannot be reconciled with machine-checkable evidence.
Verification checkpoint
Trace one released synthetic claim backwards through review, confidence, verification, evidence graph, tool result or source span to an immutable input. Then trace one rejected claim to the earliest failed contract. Re-run the package from its manifest and compare content digests. The checkpoint passes only when a second reviewer can reconstruct the accepted and blocked paths, identify every assumption and reproduce the release decision without access to the original conversation. Record unresolved risk instead of completing a missing fact with generated text.
Sources
- AI Risk Management Framework 1.0, a lifecycle framework for governing, mapping, measuring and managing context-dependent AI risk.
- Generative AI Risk Management Profile, a cross-sector profile covering risks and controls specific to generative systems.
- ISO/IEC 23894:2023 AI risk-management guidance, guidance for integrating AI-specific risk management into activities and decisions.
- ISO/IEC 25059:2023 quality model for AI systems, a quality vocabulary for specifying and evaluating AI-system characteristics.