Confidential mandate
Parallel File-System Metadata Recovery Leader
Urgent / Replacement
Parallel File-System Metadata Recovery Leader mandate in Hong Kong · Aviation Engineering Simulation
An aviation engineering group needs a twelve-month executive to restore metadata service performance and establish authoritative design-compute recovery procedures for simulation workloads on shared parallel file systems.
The mandate
A namespace reorganisation multiplied small-file and directory operations against a shared parallel file system, creating metadata service pressure while capacity and bulk bandwidth remained healthy. Simulation jobs require reliable checkpoint recovery and engineers need confidence in which design artefacts are authoritative across unmanaged working spaces.
The interim must start in Hong Kong within three weeks and serve for twelve months, with a possible three-month extension only if a major certification simulation shifts. A permanent HPC storage search begins after two aircraft programmes run through the target namespace and checkpoint pattern, expected in month six. The successor will overlap for six weeks and command the final metadata-failure exercise.
Handover requires metadata demand to remain within measured headroom during peak job launch and checkpoint, namespace and small-file patterns to have owned design limits, authoritative artefacts and permissions to reconcile, restore and failover to pass three exercises, unmanaged copies to be removed or governed, and the permanent leader to accept capacity, migration, defect and risk registers.
The interim may stop namespace changes, throttle job launches, set file and directory standards, quarantine unmanaged data, sequence migrations and approve up to HKD 140 million within the recovery plan. Permanent hiring, simulation-method changes, new aircraft programme commitments, supplier termination and decisions above HKD 30 million require council approval. The seat cannot alter design-certification evidence criteria.
Engineering model validation, solver selection, workstation services and replacement of unrelated corporate storage are explicitly out of scope. The mandate owns parallel file-system metadata and namespace design, checkpoint interaction, authoritative engineering artefacts, recovery testing, supplier remediation and the operating team supporting simulation throughput.
Why this seat is open
The disruption showed that terabytes and aggregate bandwidth concealed the metadata path governing whether simulation work could begin, checkpoint and be trusted. Executive departure left storage, workflow and design teams solving different symptoms while engineers created new copies. Temporary authority is needed to restore one evidence-bearing namespace and transfer permanent leadership after two real programmes prove it.
What you will own
- Reconstruct metadata demand across job launch, dependency scan, checkpoint, output, cleanup, permission, backup and user workaround patterns.
- Establish authoritative namespaces, ownership, retention and access controls for models, inputs, checkpoints, results and certification evidence.
- Decide sharding, metadata placement, caching, directory design, small-file aggregation, quotas and workload throttling using measured operations.
- Reconcile unmanaged copies and permission changes without overwriting more authoritative or legally retained engineering artefacts.
- Design failure tests for metadata-server loss, failover lag, journal recovery, namespace corruption, permission drift and checkpoint interruption.
- Align scheduler, workflow and file-system controls so concurrency respects metadata as well as compute and bulk bandwidth limits.
- Transfer namespace standards, measured envelopes, migration status, restore evidence, supplier defects and escalation authority through successor-led command.
Candidate qualifications
- Held executive or platform leadership for parallel file systems serving large engineering, scientific or media compute estates.
- Recovered metadata saturation or namespace failure that occurred despite available storage capacity and sequential bandwidth.
- Redesigned small-file, directory, permission and checkpoint patterns using production operation counts and workflow evidence.
- Preserved authoritative technical artefacts while removing unmanaged copies created during a severe storage event.
- Led metadata failover, journal recovery and restore exercises at representative namespace and concurrency scale.
- Handed a recovered HPC storage platform to permanent leadership through observed peak and failure operations.
More seats like this one
This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.