Confidential mandate

Retrieval-Grounded Generation Audit Director

Urgent / Unplanned

Retrieval-Grounded Generation Audit Director mandate in Madrid, Spain · Aviation Maintenance Knowledge Systems

A Madrid aviation-maintenance platform commissions a four-month audit after its assistant blended superseded instructions across aircraft effectivities, requiring traceable grounding evidence and a controlled deployment opinion.

The mandate

The client must determine whether a retrieval-grounded assistant can safely support twelve defined maintenance task families across two aircraft fleets. A quality review found answers that cited valid manuals but combined steps from superseded revisions or from aircraft with different modification status, while current dashboards credited any retrieved citation without proving that the generated instruction followed the cited passage or its effectivity conditions.

The commissioned output is a Retrieval Grounding, Revision and Effectivity Audit Book. It will contain a corpus and index lineage map, frozen expert-authored challenge set, retrieval and answer-evidence measures, citation-entailment findings, revision and tail-effectivity controls, failure taxonomy, replayable audit harness and a release, constrained-release or no-release opinion for each task family.

The four-month engagement starts 18 January 2027. Milestone one, the signed corpus population and audit protocol, is due 1 February; milestone two, the indexed-lineage analysis and locked challenge set, is due 1 March; milestone three, the blinded retrieval, generation and workflow test result, is due 5 April; milestone four, the final opinion, exception ledger and continuing control pack, is due 14 May.

Acceptance requires Technical Publications to trace every sampled chunk to an approved document revision, Maintenance Engineering to reproduce effectivity judgements without seeing the assistant’s answer, and Quality to rerun the complete critical-case harness from frozen components. Every safety-significant unsupported statement, mixed revision and wrong-tail instruction must be dispositioned; both sponsors must approve task boundaries, score thresholds and residual exceptions before the audit closes.

The client will provide controlled manuals, revision histories, engineering orders, modification and tail-effectivity records, index snapshots, chunking and ranking configurations, model and prompt versions, user query logs and representative workflow access. Licensed engineers in Madrid, Toulouse and Lisbon will author and blind-review cases; the consultant will not amend maintenance instructions, approve aircraft release, choose the production model vendor or audit unrelated knowledge applications.

Why this is external work

The product team built both the retrieval pipeline and the dashboard that currently labels citations as grounded evidence. Technical-publication specialists know revision and effectivity rules but have not audited semantic retrieval or generated-answer entailment at scale. External direction provides methodological independence and enough aviation context to prevent a fluent, well-cited response from being mistaken for an authorised instruction.

What you will own

  • Reconcile manuals, revisions, engineering orders, effectivity tables, index snapshots and chunk transformations into a lineage graph that exposes stale or unauthorised retrieval content.
  • Design a blinded challenge set spanning fleet, tail, modification, task phase, supersession, exception and intentionally insufficient evidence across the twelve approved families.
  • Measure retrieval completeness and contamination separately from answer support, citation correctness and faithful expression of conditions or prohibitions.
  • Test multi-document answers for mixed revisions, incompatible aircraft effectivity, omitted cautions, fabricated sequencing and conclusions unsupported by any retrieved source.
  • Observe engineers using the assistant under realistic time and interface conditions, capturing automation bias, verification behaviour and unsafe workaround patterns.
  • Set task-specific release gates, mandatory source views, abstention rules and index-freshness checks tied to the consequence of a grounding failure.
  • Transfer the replayable harness, source fingerprints, expert rationales and exception workflow so Quality can audit every later corpus, retriever or model change.

Candidate qualifications

  • Led an independent retrieval-augmented generation audit in aviation, engineering, life sciences or another domain governed by revision-controlled technical evidence.
  • Detected answers that carried plausible citations yet contradicted source meaning, mixed document states or ignored asset-specific applicability.
  • Built evaluation sets with expert-blinded answers, hard negatives, insufficient-evidence cases and traceable rationales rather than synthetic question generation alone.
  • Audited corpus ingestion, chunking, ranking, prompting and model output as one evidence chain across version changes.
  • Converted grounding results into task-level release restrictions, human-verification controls and continuing regression tests that operators could apply.
  • Presented adverse findings to technical, safety and product executives when a commercially important assistant could not support its claimed operating boundary.

Non-negotiables

  • The named director must lead audit design and final opinions, attend the Madrid reviews and complete the Toulouse and Lisbon workflow observations.
  • Holds demonstrable experience with revision, configuration or effectivity-controlled technical content; general-purpose chatbot evaluation is insufficient.
  • No licensed manual content, engineering order, aircraft record or user query may be copied outside the protected client environment.
  • Will return a constrained or no-release opinion for any task family whose source authority and effectivity cannot be reproduced reliably.
  1. 49 words maximum. Describe a cited generated answer you found materially unsupported because of revision, configuration or effectivity, and the control imposed.
  2. 49 words maximum. How would you separate retrieval failure from citation-entailment and workflow-verification failure during the first audit month?
  3. 49 words maximum. Confirm four-month capacity, Madrid, Toulouse and Lisbon attendance, and any relevant model or retrieval-vendor relationship.

This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.