Confidential mandate

AI Accelerator Compiler Performance Leader

Urgent / Replacement

AI Accelerator Compiler Performance Leader mandate in Taipei, Taiwan · AI Accelerator Semiconductors

After its compiler chief left during silicon qualification, an accelerator company needs an eighteen-month executive to recover model performance, stabilise release evidence and transfer a durable software organisation.

The mandate

The compiler chief resigned after a silicon revision and framework update invalidated headline throughput results three weeks before qualification. Kernel, graph and runtime teams can repair isolated regressions, but no executive now owns numerical equivalence, representative model coverage, host-device interactions and the decision to ship software against contracted customer workloads.

The interim must establish a Taipei base within three weeks and lead for eighteen months. A permanent search opens after the second silicon revision clears a shared model suite, expected near month twelve, with eight weeks reserved for the successor to chair performance allocation and approve a compiler-runtime release before taking accountability.

Handover requires three representative model families to meet signed accuracy, latency, throughput, memory and power corridors across both supported silicon revisions; every benchmark must have reproducible provenance; two rollback and compatibility drills must pass; and the successor must accept the unresolved framework, kernel and architecture dependency ledger. Peak throughput on selected graphs is insufficient.

The interim may stop software releases, reallocate engineering and laboratory capacity, set benchmark and numerical standards, appoint temporary workstream leads and move up to TWD 480 million within the approved portfolio. Silicon architecture changes, permanent executive hiring, customer contract concessions, open-source licensing shifts and decisions above TWD 140 million require council approval; performance claims cannot omit accuracy or host cost.

Chip physical design, fabrication yield, commercial pricing and customer model ownership are explicitly out of scope. The seat owns compiler, kernel and runtime performance, framework integration, reproducibility, release evidence and engineering leadership while hardware and customer product authorities retain their respective decisions.

Why this seat is open

The qualification regression exposed a software organisation optimising component benchmarks without one owner for customer-visible system performance. The chief’s abrupt departure left a release decision suspended between silicon and framework teams. Temporary executive authority is needed to recover evidence and organisation before long-term leadership steers the next architecture generation.

What you will own

  • Reconstruct invalidated results across model, framework, graph transformations, kernels, runtime, driver, host, silicon revision and measurement configuration.
  • Define representative language, vision and recommendation workload suites with accuracy, shape, sparsity, memory, latency, throughput, power and host-overhead corridors.
  • Decide optimisation priorities across graph lowering, fusion, scheduling, quantisation, kernels, communication and runtime using customer value and architectural learning.
  • Establish numerical equivalence and tolerance controls that expose performance gained through unsupported precision, operator substitution or altered model semantics.
  • Install release gates for framework compatibility, deterministic build, benchmark provenance, regression, fallback and customer reproduction.
  • Negotiate workload evidence and qualification expectations with framework partners and strategic customers without accepting curated demo conditions.
  • Transfer the performance atlas, release history, architecture feedback, customer exceptions and talent map through a successor-led qualification cycle.

Candidate qualifications

  • Led compiler and runtime engineering for AI accelerators, GPUs or specialised compute used by production model workloads.
  • Recovered a system-performance regression spanning framework, compiler, kernel, driver and silicon rather than one isolated optimisation.
  • Governed numerical equivalence across precision, quantisation and operator changes while measuring accuracy and customer semantics.
  • Built representative benchmark suites whose configurations, builds, datasets and hardware states were independently reproducible.
  • Negotiated qualification evidence with major framework, cloud or model-platform engineering teams under launch pressure.
  • Handed a semiconductor software organisation to permanent leadership after stable multi-revision releases and measured customer reproduction.

Non-negotiables

  • Can start within three weeks, work on site in Taipei, attend weekly Hsinchu laboratories and travel quarterly to a partner or customer hub.
  • Will block performance claims or releases that lack numerical, configuration and host-cost evidence despite silicon launch commitments.
  • Brings direct AI compiler and runtime leadership; kernel specialism, chip architecture or generic developer tooling alone is insufficient.
  • Must disclose current accelerator, foundry, framework, cloud and compiler-vendor relationships before qualification evidence is shared.
  1. 49 words maximum. State your earliest Taipei start and any obligation incompatible with eighteen months of executive authority.
  2. 49 words maximum. Describe a compiler performance gain you rejected because numerical or system evidence made it invalid.
  3. 49 words maximum. Which cross-revision regression would make you stop a release despite meeting one headline benchmark?

This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.