Confidential mandate

Multilingual Contact-Centre AI Evaluation Architect

Planned Hiring / New

Multilingual Contact-Centre AI Evaluation Architect mandate in Kuala Lumpur, Malaysia · Regional Aviation Technology

A regional airline commissions a four-month evaluation system for multilingual service AI, linking generated answers and summaries to booking state, disruption recovery, customer outcomes and agent workload.

The mandate

The airline is piloting generated agent assistance, call summaries and self-service answers across Bahasa Malaysia, English, Thai and Mandarin, but present evaluation separates language quality from booking and disruption outcomes. The defined problem is to establish whether assistance remains grounded in live journey state and actually improves safe resolution under routine and disrupted operations.

The deliverable is a Multilingual Service AI Evaluation System containing a task taxonomy, language and channel sample, booking-state fixtures, retrieval and tool lineage, fluent-review rubric, disruption scenarios, summary fidelity tests, service-outcome measures, agent-workload analysis, release gates and a maintained regression suite. It must identify language-task cells unsuitable for automation.

Work begins 11 January 2027. Milestone one at month one is the signed task, language and evidence design; milestone two at month three provides baseline findings, witnessed disruption tests and urgent containment recommendations; milestone three at month four delivers the accepted system, reproducible casebook, release matrix and transfer to service-quality owners.

Acceptance requires language and operations teams to reproduce sampled outputs from customer context, booking state, source content, model and prompt versions and tool results. Each approved cell must meet signed grounding, action, summary and escalation corridors during nominal and disrupted journeys; average agent handling time or customer sentiment alone cannot establish acceptance.

The client will provide consented interactions, booking and disruption fixtures, knowledge versions, model endpoints, tool logs, agent workflows, quality reviews and outcome measures in a secure environment. Language leads will resolve reference disagreement; the consultant will not contact passengers, change fare or reaccommodation policy, deploy models or operate quality controls after handover.

Why this is external work

Customer operations sponsors productivity while AI teams report model-level quality and country teams use different review standards. No internal owner joins language evidence to live booking state and disruption consequence. An external architect can build one challengeable system without defending the current pilot or owning its eventual rollout.

What you will own

  • Define service tasks by language, channel, booking state, disruption, customer vulnerability, action consequence and required human authority.
  • Build controlled fixtures for schedule change, cancellation, missed connection, baggage, refund, loyalty and ambiguous itinerary scenarios.
  • Trace generated answers and summaries to knowledge, passenger context, model, prompt, tool and system-of-record evidence.
  • Establish fluent review for factuality, clarity, tone, omission, unsupported commitment, code switching and effective escalation.
  • Compare agent assistance and self-service through resolution, recontact, correction, transfer, handling effort and customer consequence.
  • Run disruption simulations under rapidly changing status, unavailable options, policy conflict and high queue pressure.
  • Transfer the regression suite, annotated casebook, decision thresholds, reviewer calibration and model-change triggers to internal owners.

Candidate qualifications

  • Built multilingual AI evaluation for aviation, travel, banking, telecommunications or another complex service environment.
  • Connected generated answer or summary quality to system state, agent action and eventual customer resolution.
  • Designed fluent-review operations across at least three relevant Asian languages with measured reviewer disagreement.
  • Tested AI assistance during disruption or rapidly changing transactional context rather than static frequently asked questions.
  • Identified a productivity gain that concealed increased correction, recontact or harmful customer commitment.
  • Delivered a reusable evaluation suite that internal service and technology teams recalibrated after engagement acceptance.

Non-negotiables

  • The named architect must attend all three Kuala Lumpur or regional operations-floor and disruption-simulation sessions.
  • No passenger record, recording, booking fixture or model output may leave the airline’s controlled environment.
  • Will disclose airline, contact-centre, model, translation and customer-service platform relationships before evidence access.
  • Must identify prohibited automation cells where grounding or escalation evidence remains insufficient despite efficiency gains.
  1. 49 words maximum. Describe one multilingual service-AI error whose operational consequence was invisible in a language-quality score.
  2. 49 words maximum. How would you test generated reaccommodation guidance while flight and booking state changes during the interaction?
  3. 49 words maximum. Confirm four-month capacity, regional travel and every airline or service-AI relationship relevant to independence.

This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.