Confidential mandate
Parallel File-System Metadata Recovery Leader
Urgent / Replacement
Parallel File-System Metadata Recovery Leader mandate in Hong Kong · Aviation Engineering Simulation
An aviation engineering group needs a twelve-month executive after namespace reorganisation saturated metadata services, stranded simulation checkpoints and undermined programme confidence in authoritative design-compute recovery.
The mandate
A namespace reorganisation multiplied small-file and directory operations against a shared parallel file system, saturating metadata servers while capacity and bulk bandwidth appeared healthy. Simulation jobs stalled, checkpoints timed out and engineers copied working sets into unmanaged spaces. The HPC storage executive departed after recovery attempts introduced inconsistent permissions and nobody could prove which design artefacts remained authoritative.
The interim must start in Hong Kong within three weeks and serve for twelve months, with a possible three-month extension only if a major certification simulation shifts. A permanent HPC storage search begins after two aircraft programmes run through the target namespace and checkpoint pattern, expected in month six. The successor will overlap for six weeks and command the final metadata-failure exercise.
Handover requires metadata demand to remain within measured headroom during peak job launch and checkpoint, namespace and small-file patterns to have owned design limits, authoritative artefacts and permissions to reconcile, restore and failover to pass three exercises, unmanaged copies to be removed or governed, and the permanent leader to accept capacity, migration, defect and risk registers.
The interim may stop namespace changes, throttle job launches, set file and directory standards, quarantine unmanaged data, sequence migrations and approve up to HKD 140 million within the recovery plan. Permanent hiring, simulation-method changes, new aircraft programme commitments, supplier termination and decisions above HKD 30 million require council approval. The seat cannot alter design-certification evidence criteria.
Engineering model validation, solver selection, workstation services and replacement of unrelated corporate storage are explicitly out of scope. The mandate owns parallel file-system metadata and namespace design, checkpoint interaction, authoritative engineering artefacts, recovery testing, supplier remediation and the operating team supporting simulation throughput.
Why this seat is open
The disruption showed that terabytes and aggregate bandwidth concealed the metadata path governing whether simulation work could begin, checkpoint and be trusted. Executive departure left storage, workflow and design teams solving different symptoms while engineers created new copies. Temporary authority is needed to restore one evidence-bearing namespace and transfer permanent leadership after two real programmes prove it.
What you will own
- Reconstruct metadata demand across job launch, dependency scan, checkpoint, output, cleanup, permission, backup and user workaround patterns.
- Establish authoritative namespaces, ownership, retention and access controls for models, inputs, checkpoints, results and certification evidence.
- Decide sharding, metadata placement, caching, directory design, small-file aggregation, quotas and workload throttling using measured operations.
- Reconcile unmanaged copies and permission changes without overwriting more authoritative or legally retained engineering artefacts.
- Design failure tests for metadata-server loss, failover lag, journal recovery, namespace corruption, permission drift and checkpoint interruption.
- Align scheduler, workflow and file-system controls so concurrency respects metadata as well as compute and bulk bandwidth limits.
- Transfer namespace standards, measured envelopes, migration status, restore evidence, supplier defects and escalation authority through successor-led command.
Candidate qualifications
- Held executive or platform leadership for parallel file systems serving large engineering, scientific or media compute estates.
- Recovered metadata saturation or namespace failure that occurred despite available storage capacity and sequential bandwidth.
- Redesigned small-file, directory, permission and checkpoint patterns using production operation counts and workflow evidence.
- Preserved authoritative technical artefacts while removing unmanaged copies created during a severe storage event.
- Led metadata failover, journal recovery and restore exercises at representative namespace and concurrency scale.
- Handed a recovered HPC storage platform to permanent leadership through observed peak and failure operations.
Non-negotiables
- Can start on site in Hong Kong within three weeks and travel monthly to engineering or infrastructure locations.
- Will accept exclusive executive responsibility and continuous escalation availability for simulation-storage events.
- Brings direct parallel metadata recovery; generic enterprise storage or capacity procurement alone is insufficient.
- Must disclose relationships with file-system, storage, scheduler, network and engineering-compute suppliers.
- 49 words maximum. State your earliest Hong Kong start and the largest metadata-operation peak you have personally recovered.
- 49 words maximum. Describe a file system with ample capacity and bandwidth that still failed its simulation users.
- 49 words maximum. Which evidence distinguishes an authoritative design checkpoint from a convenient unmanaged copy?
This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.