Confidential mandate

Low-Resource Speech AI Evaluation Director

Planned Hiring / New

Low-Resource Speech AI Evaluation Director mandate in Lagos, Nigeria · Multilingual Communications Platforms

A West African communications platform commissions a six-month evaluation system for low-resource speech models, covering language variation, acoustic reality, task usefulness and equitable service performance across markets.

The mandate

The platform plans to automate service transcription and routing across several West African languages, yet its reported word-error improvements come from small studio-heavy datasets and do not reveal task failure, code switching or acoustic inequality. The defined problem is to create a repeatable evidence system that supports language-by-use decisions rather than one claim of regional readiness.

The deliverable is a Low-Resource Speech Evaluation Observatory containing language and task taxonomies, recording provenance, speaker and acoustic sampling rules, transcription-quality controls, recognition and intent measures, code-switch tests, subgroup uncertainty, service-outcome links, model-change gates and a maintained evidence repository. It must permit a “not ready” conclusion for individual cells.

The engagement begins 11 January 2027. Milestone one at month one is the signed sampling and consent protocol; milestone two at month three supplies baseline evaluations and data-gap findings; milestone three at month five provides prospective service pilots and error analysis; milestone four at month six delivers the accepted observatory, rerun evidence and operating transfer.

Acceptance requires the language-operations team to reproduce sampled scores from retained audio, transcripts, adjudications and model versions, and product quality to trace priority errors to routing or service outcomes. Confidence and reviewer disagreement must be reported for sparse cells; pooled regional averages, translated studio prompts or unlabeled code switching will prevent final acceptance.

The client will provide consented recordings, service tasks, language-partner access, telephony channels, model endpoints, prior transcripts, routing outcomes and secure regional storage. Community and language leads will decide acceptable collection and use; the consultant will not scrape audio, claim speaker representation, retrain production models or operate evaluation after transfer.

Why this is external work

Speech engineers are rewarded for aggregate model improvement, while country teams see failures without shared measurement or enough evaluation capacity. Internal quality methods were designed for high-resource languages and stable contact-centre audio. External specialist delivery can build evidence around sparse, varied conditions without making commercial launch pressure the definition of readiness.

What you will own

  • Define language, dialect, code-switch, task, channel, device, noise, bandwidth and speaker cells relevant to intended service use.
  • Design consented sampling and provenance controls that preserve contributor permissions, recording context, transcription history and permitted reuse.
  • Establish fluent transcription, adjudication and reviewer-agreement methods with uncertainty reported where reference labels remain contested.
  • Measure recognition, intent, entity and routing performance alongside service completion, escalation burden and harmful misunderstanding.
  • Diagnose phonetic, lexical, acoustic, channel and code-switch failures without treating country or language labels as homogeneous cohorts.
  • Set release cells and minimum evidence for shadow use, assisted operation, broader automation or explicit exclusion from deployment.
  • Transfer datasets, sampling logic, evaluation code, adjudication records, dashboards and model-change triggers to regional owners.

Candidate qualifications

  • Built speech-recognition or spoken-language evaluation for low-resource, multilingual or code-switched production settings.
  • Designed consented audio and transcription operations where speaker coverage, dialect boundaries and reference disagreement were material.
  • Connected word or character error to task, routing and customer outcomes rather than presenting one technical score.
  • Quantified uncertainty for sparse language cells and resisted pooling that concealed a materially weaker cohort.
  • Worked with local language authorities, annotators and service teams without treating translated high-resource prompts as representative speech.
  • Delivered an evaluation platform that regional product teams could refresh when models, channels or intended tasks changed.

Non-negotiables

  • The named director must complete monthly Lagos laboratory weeks and attend all three scheduled in-market recording or workflow reviews.
  • No audio, transcript, speaker attribute or model output may leave approved regional storage or enter personal AI services.
  • Will disclose speech-provider, annotation, telecom and language-data relationships before accessing protected recordings.
  • Must state where evidence is too sparse for launch, even if doing so reduces the announced market or language count.
  1. 49 words maximum. Describe a low-resource speech cohort whose task failure was hidden by aggregate word-error improvement.
  2. 49 words maximum. How would you measure code-switched service performance when fluent reviewers disagree on a reference transcript?
  3. 49 words maximum. Confirm six-month capacity, West African travel and every speech-data or model-provider relationship relevant to independence.

This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.