Confidential mandate

Accelerated-Compute Queue Economics Director

Planned Hiring / New

Accelerated-Compute Queue Economics Director mandate in Boston, United States · Computational Drug Discovery

A computational-discovery company needs an independent four-month decision pack after GPU expenditure doubled while high-value molecular programmes continued waiting behind poorly classified and weakly governed research workloads.

The mandate

GPU invoices have doubled in nine months, yet experiment leads report longer time-to-first-run and repeatedly reserve scarce accelerators through oversized queues. Finance sees utilisation above eighty per cent, scientists see blocked programmes, and platform teams cannot distinguish productive training, interactive exploration, failed jobs, checkpoint recovery or capacity held idle by workflow design. The defined problem is to establish which scheduling and demand choices destroy scientific throughput per unit of compute.

The named deliverable is an Accelerated-Compute Scheduling and Unit-Economics Decision Pack. It must contain a workload taxonomy, reconciled cost and useful-work baseline, queue-policy simulator, capacity alternatives, chargeback design, target scheduler controls, migration sequence and a board-ready capital recommendation. It is a decision artefact, not an implementation programme or a generic cloud-cost review.

Four milestones govern the work: by day fifteen, sign off the inventory and invoice reconciliation; by week seven, deliver queue traces, failure waste and programme-delay economics; by week twelve, complete controlled scheduling experiments and capacity scenarios; and by week sixteen, submit the final policy, investment case and ninety-day implementation backlog. Each milestone has its own review and invoice.

Acceptance rests with the chief scientific officer and finance sponsor. They will accept the pack only if at least ninety-five per cent of accelerator spend reconciles to a workload class, proposed priorities replay against twelve months of queue history, scientific delay is valued separately from nominal utilisation, and every recommended control has an accountable owner, measurable threshold and reversible introduction path.

The client will provide scheduler events, cloud and colocation invoices, experiment metadata, failure logs, research milestones, procurement terms, researcher interviews and a data engineer for secure extraction. The consultant will not redesign molecular models, select scientific targets, negotiate hardware purchases or operate the production cluster; inaccessible or legally restricted datasets will be recorded as evidence limits rather than estimated away.

Why this is external work

Internal platform staff own the scheduler configuration whose effects must be independently examined, while research teams have strong incentives to defend priority access. Finance can price invoices but cannot value delayed scientific learning or distinguish useful accelerator time from apparently busy waste. A specialist external view gives the steering group a neutral basis for deciding policy before it commits the next capacity tranche.

What you will own

  • Reconcile reserved, on-demand, colocation and owned accelerator costs to jobs, users, programmes, failures, waiting time and stranded allocations.
  • Classify workloads by scientific decision value, accelerator topology, memory need, runtime uncertainty, checkpoint behaviour, pre-emption tolerance and data locality.
  • Build a replayable queue model that compares fair-share, deadline, gang, backfill and quota policies against actual programme milestones.
  • Quantify useful tokens, conformations, simulations or validated experiments per fully burdened compute unit instead of relying on device utilisation alone.
  • Test scheduling interventions on a controlled workload slice, capturing starvation, failed recovery, researcher behaviour and cross-cluster transfer costs.
  • Price capacity choices spanning reservations, elastic cloud, dedicated clusters, brokered supply and workload deferral under demand and failure scenarios.
  • Deliver the signed decision pack, implementation backlog, measurement dictionary and unresolved-evidence register at the fourth milestone.

Candidate qualifications

  • Designed or materially changed GPU or HPC scheduling for research workloads with heterogeneous runtimes, topology constraints and scientific priorities.
  • Reconciled accelerator telemetry to invoices and business or research outcomes, exposing waste that high aggregate utilisation had concealed.
  • Built queue simulations from production traces and used them to test pre-emption, backfill, quotas, reservations or deadline-aware scheduling.
  • Evaluated owned, colocated and elastic compute economics including network, storage, failure recovery, software and specialist operating labour.
  • Worked credibly with scientists who resist workload classification and can separate legitimate exploratory freedom from unmanaged capacity capture.
  • Produced a capital or operating decision paper accepted jointly by technical, scientific and finance executives without implementing the recommendation yourself.

Non-negotiables

  • Can work in Boston and San Diego during the specified research-hub weeks while handling restricted experiment metadata under client controls.
  • Has personally analysed production scheduler traces; procurement benchmarking or dashboard-only FinOps experience does not meet the requirement.
  • Will remain independent of accelerator manufacturers, cloud resellers, schedulers and capacity brokers throughout the engagement.
  • Accepts that the fee closes on accepted artefacts and evidence tests, not hours expended or a promised utilisation percentage.
  1. 49 words maximum. Name the scheduler traces you would request first to distinguish busy accelerators from scientifically useful work.
  2. 49 words maximum. Describe one queue-policy change that improved utilisation but worsened the outcome the institution actually valued.
  3. 49 words maximum. How would you price a day of delayed molecular learning without inventing false precision?

This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.