Confidential mandate

Real-Time AI Inference Fleet Recovery Leader

Urgent / Replacement

Real-Time AI Inference Fleet Recovery Leader mandate in Singapore · Live Commerce and Digital Services

After a peak-event collapse exposed hidden accelerator contention, a consumer platform needs a nine-month executive to recover real-time inference reliability, reset cost discipline and hand over an exercised regional fleet.

The mandate

The inference-platform executive was dismissed after a major live-commerce event saturated shared accelerator memory, sending recommendation and moderation models into competing fallback paths for seventy-six minutes. Customer journeys recovered, but engineers cannot reconcile reserved capacity, scheduler priority, batching policy and model-specific latency, while product teams continue to launch endpoints against incomparable service promises.

An experienced operator must assume the Singapore decision seat within twelve days for nine months. The permanent search begins after the first successful regional peak rehearsal in month four, with five weeks reserved for the chosen successor to shadow capacity decisions and then chair a live traffic review before taking accountability.

Handover is earned when the twelve most consequential inference journeys remain inside agreed tail-latency and availability budgets through two peak windows, unit cost stays within its approved corridor for three billing cycles, every production model has exercised canary and rollback paths, and the successor has commanded a multi-region accelerator-loss scenario. Higher average utilisation alone does not close the work.

The interim may set admission and workload-placement rules, freeze model promotions, reassign reserved capacity, alter on-call design, appoint temporary recovery leads and move up to SGD 12 million within the authorised platform envelope. Multi-year hardware commitments above SGD 8 million, permanent appointments, workforce restructuring and changes to customer service promises require executive approval; the role cannot lower moderation or ranking safeguards to improve speed.

Base-model training, recommendation objectives, advertising auction logic and the corporate network refresh remain outside scope. This mandate is confined to the real-time serving fleet, its capacity and cost controls, model-release interfaces and regional command system, although upstream teams must supply measurements needed to operate those boundaries.

Why this seat is open

The event review found that a nominal capacity shortage was amplified by policy conflicts no single leader was authorised to resolve. Removing the incumbent created a safer governance reset but left peak-season decisions split among platform, product and finance. The company needs temporary executive command until a permanent operator can inherit tested rules rather than undocumented compromises.

What you will own

  • Reconstruct the peak failure from request class through routing, batching, memory allocation, model execution, fallback and customer-visible consequence, preserving disputed evidence.
  • Set service envelopes for each critical inference journey covering tail latency, throughput, accuracy guardrails, availability, degradation behaviour and maximum unit cost.
  • Decide placement across reserved accelerators, elastic cloud capacity, specialised inference services and CPU fallback using workload evidence and failure isolation.
  • Install a promotion gate that requires representative load tests, cost forecasts, observability, safe degradation, canary criteria and rehearsed rollback for every material model change.
  • Rebalance scheduler priority and admission control so safety, moderation and high-consequence transactional models cannot be starved by commercially promoted traffic.
  • Negotiate regional capacity reservations and supplier remedies against demand scenarios, hardware obsolescence, deployment lead time and interruptible-service risk.
  • Transfer the fleet topology, economic baseline, release decisions, incident library, capacity calendar and talent assessment through a successor-led peak exercise.

Candidate qualifications

  • Held enterprise authority for a large low-latency inference, recommendation, search or advertising-serving platform operating across several Asian markets.
  • Personally resolved accelerator memory, batching, scheduler or routing contention during a customer-visible traffic event and can evidence the resulting service decision.
  • Built model-specific latency and cost envelopes that reconciled telemetry, reserved capacity, fallback behaviour and product service commitments.
  • Governed heterogeneous GPU or accelerator placement across owned, reserved and elastic environments without masking idle commitments as consumed workload.
  • Introduced release controls that joined model-quality thresholds with load, observability, canary and rollback evidence in a continuously changing fleet.
  • Handed a high-availability platform to permanent leadership after exercising regional failover, capacity loss and cross-functional incident command.

Non-negotiables

  • Available in Singapore within twelve days for an exclusive assignment and able to join the regional severity-one command roster immediately.
  • Will work on site through the first three months and travel monthly to a designated engineering or high-volume market hub.
  • Has directly governed production inference economics and tail latency; training infrastructure or generic SRE leadership alone will not qualify.
  • Must disclose current cloud, accelerator, managed-inference and regional consumer-platform relationships before commercial information is shared.
  1. 49 words maximum. State your earliest Singapore start date and the commitment you would need to leave before holding exclusive fleet authority.
  2. 49 words maximum. Describe one inference contention event you commanded, the scheduler or placement decision you made and its measured peak result.
  3. 49 words maximum. Which tail-latency and cost evidence would make you block a commercially important model promotion?

This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.