Confidential mandate

On-Device Generative AI Thermal and Latency Director

Planned Hiring / New

On-Device Generative AI Thermal and Latency Director mandate in Seoul, South Korea · Premium Consumer Devices

A premium-device maker commissions a twelve-week optimisation engagement to determine which generative workloads can run locally within measurable latency, battery, thermal, memory and output-quality limits across its next hardware family.

The mandate

Prototype teams can make summarisation, image editing and assistant features run on individual devices, but each demonstration uses a different model variant, context size, ambient condition and quality measure. The bounded problem is to establish a reproducible deployment envelope for three hardware classes and decide which workloads remain local, become hybrid, or should not ship under the next release’s constraints.

The required artefact is an On-Device Generative AI Performance Book containing reference workloads, representative prompt distributions, model and compiler configurations, quantisation evidence, memory and power traces, heat-soak behaviour, quality-loss tolerances, fallback rules and an engineering decision matrix. It must expose interactions hidden by a single latency or benchmark score.

Work begins on 4 January 2027. Milestone one, due 22 January, is the signed workload corpus, measurement protocol and instrument reconciliation; milestone two on 19 February delivers the comparative model-compiler experiments and failure analysis; milestone three on 26 March supplies the accepted deployment envelopes, optimisation backlog, product recommendations and reproducible benchmark repository.

The Head of Device Platforms and product-quality director will accept the engagement only when their engineers can rerun the agreed workloads on each reference device within signed measurement tolerance, reproduce the quality deltas and trace every recommendation to thermal, power, memory and user-experience evidence. A peak demo, laboratory-only ambient condition or vendor estimate will not pass acceptance.

The client will supply model checkpoints, compiler access, prototype devices, battery and thermal instrumentation, representative product tasks, baseline quality suites and named silicon and software engineers. The consultant may instrument approved development builds but will not choose final consumer features, redesign the system-on-chip, certify battery safety or own implementation after the accepted optimisation backlog transfers.

Why this is external work

Model, silicon and product teams each optimise the metric attached to their delivery gate, so no internal group owns the trade-off across all device constraints. The portfolio also lacks a common method for comparing local and hybrid execution as models and compilers change. External delivery provides neutral measurement discipline before launch choices harden into marketing promises.

What you will own

  • Define representative text, image and multimodal workloads with context, output, concurrency, ambient and interaction conditions tied to intended product use.
  • Reconcile power, thermal, memory, latency and utilisation instrumentation across prototype devices so platform comparisons share one defensible measurement basis.
  • Benchmark model size, quantisation, pruning, speculative execution, cache policy and compiler alternatives while preserving signed output-quality tolerances.
  • Identify sustained heat, battery drain, memory pressure and throttling breakpoints that short demonstrations or room-temperature averages conceal.
  • Compare local, hybrid and remote execution for privacy, responsiveness, offline continuity, quality, energy and service-cost consequences.
  • Rank engineering interventions by user value, technical confidence, hardware reach, implementation effort and risk of invalidating other device workloads.
  • Deliver the versioned benchmark repository, chosen envelopes, rejected configurations, residual uncertainties and refresh instructions to the internal performance team.

Candidate qualifications

  • Led on-device deployment of generative, language, vision or multimodal models across more than one production silicon or device class.
  • Optimised models through quantisation, compilation, memory management or runtime scheduling and measured the resulting quality loss on representative tasks.
  • Instrumented sustained thermal, battery, latency and memory behaviour on consumer hardware rather than relying on simulator or accelerator specifications.
  • Made local-versus-hybrid execution decisions that balanced privacy and user experience against device constraints and cloud cost.
  • Built a performance suite that remained reproducible across model, compiler, operating-system and prototype-hardware changes.
  • Presented a device feature recommendation that executives altered because measured heat, battery or quality evidence contradicted a successful demonstration.

Non-negotiables

  • The engagement lead must work weekly in the Suwon laboratory and attend both Gumi device-line sessions while remaining available to Seoul sponsors.
  • No prototype, checkpoint, compiler binary, prompt corpus or measurement trace may leave the client-controlled engineering environment.
  • Commercial relationships with silicon, model-compression, compiler, battery or benchmark vendors under comparison must be declared before testing.
  • Will reject any configuration whose sustained evidence fails the signed envelope, regardless of launch-event or feature-marketing commitments.
  1. 49 words maximum. Describe an on-device generative workload you optimised and the quality sacrifice required to meet its sustained thermal envelope.
  2. 49 words maximum. How would you detect a benchmark that understates memory pressure or battery drain during real interactive use?
  3. 49 words maximum. Confirm twelve-week capacity, Korean laboratory travel and all supplier relationships relevant to model or compiler selection.

This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.