Confidential mandate

Pre-emptible HPC Checkpoint Recovery Architect

Planned Hiring / New

Pre-emptible HPC Checkpoint Recovery Architect mandate in Zurich, Switzerland · Catastrophe Risk Research

A catastrophe-risk institute needs an eight-week recovery blueprint after interrupted simulations repeatedly restarted from unusable checkpoints, consuming critical forecast windows and scarce national supercomputing allocations.

The mandate

Long-running catastrophe simulations have been moved onto pre-emptible capacity, but application-level checkpoints sometimes reference incomplete parallel files, incompatible binaries or external workflow state that cannot be replayed. Jobs appear checkpointed until a weather-driven production deadline demands restart, when teams discover silent corruption or days of recomputation. The defined problem is to establish trustworthy recovery semantics across model, scheduler, storage and workflow layers.

The named deliverable is a Pre-emptible HPC Checkpoint Recovery Blueprint with a tested reference workflow. It must include workload classes, state boundaries, consistency protocol, storage and metadata design, scheduler integration, version compatibility, integrity evidence, restart objectives, cost model, failure taxonomy and implementation backlog for the two highest-value simulation families.

Three milestones fit eight weeks: at the end of week two, approve the workload and failure baseline; by week five, demonstrate injected interruption and recovery experiments across representative node and storage faults; and by week eight, deliver the blueprint, reference workflow, runbook, economic threshold and accepted adoption plan. Each output gates its milestone invoice.

The research director and model production lead will accept the work only if recovered runs reproduce agreed scientific invariants, checkpoints prove complete before jobs relinquish capacity, incompatible code or data versions fail safely, restart time is measured at production scale, and client engineers can run the interruption suite without the consultant. Bitwise identity is required only where the scientific method specifies it.

The client will provide representative models, schedulers, parallel file systems, object storage, workflow metadata, container and library versions, allocation prices, failure records and cleared engineers. The consultant will not modify scientific equations, validate catastrophe assumptions, procure compute, operate production forecasts or redesign unrelated research workflows; restricted models may be represented by approved test harnesses.

Why this is external work

Model teams know their science, storage teams know infrastructure and scheduler teams know allocation, but no internal owner governs a checkpoint as a cross-layer recovery contract. Past fixes optimised individual components and were declared successful before production restart. Independent HPC resilience expertise is needed quickly because the next seasonal modelling window cannot be used as another experiment.

What you will own

  • Classify simulation state across memory, accelerator, parallel files, object stores, workflow engines, external datasets and downstream production steps.
  • Define atomicity, consistency, integrity, version and completion evidence for each checkpoint family and permitted restart boundary.
  • Design interruption experiments covering node loss, scheduler pre-emption, partial write, metadata delay, storage unavailability and changed software environments.
  • Measure checkpoint overhead, useful interval, restart duration, recomputation, storage footprint and scientific validation under representative scale.
  • Establish when application, system-level, incremental or workflow checkpoints are justified by failure exposure and capacity economics.
  • Demonstrate a reference workflow that refuses incomplete state, preserves provenance and resumes without duplicating downstream effects.
  • Deliver the tested blueprint, scripts, runbook, cost thresholds, adoption sequence and unresolved scientific constraints.

Candidate qualifications

  • Designed checkpoint and restart mechanisms for large parallel CPU or accelerator workloads operating under pre-emption or frequent failure.
  • Diagnosed recovery failures spanning application state, parallel storage, workflow metadata, software version and scheduler behaviour.
  • Defined scientific equivalence or reproducibility tests appropriate to numerical simulation rather than assuming bitwise identity everywhere.
  • Ran production-scale fault injection that exposed checkpoint completeness or restart-time assumptions missed by component tests.
  • Modelled checkpoint interval and storage choices against failure probability, recomputation, allocation and deadline consequence.
  • Left an HPC team with executable recovery tests and operating thresholds it could maintain independently.

Non-negotiables

  • Can complete the supercomputing-centre visit and two production workshops inside the non-extendable eight-week schedule.
  • Will protect restricted model and hazard data and disclose relationships with compute, storage and orchestration vendors.
  • Brings hands-on parallel checkpoint recovery; backup strategy or generic disaster recovery alone is insufficient.
  • Accepts client-run interruption tests and scientific invariants as conditions for final acceptance.
  1. 49 words maximum. Describe a checkpoint that was complete at the application layer but unrecoverable as a production workflow.
  2. 49 words maximum. Which injected storage fault most efficiently exposes false checkpoint atomicity in a parallel simulation?
  3. 49 words maximum. How would you decide whether checkpoint overhead is justified for a pre-emptible workload?

This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.