Confidential mandate
Critical-Infrastructure Cloud Resilience Leader
Urgent / Replacement
Critical-Infrastructure Cloud Resilience Leader mandate in Sydney, Australia · Electricity Market Infrastructure
Following a control-plane outage and executive resignation, an electricity-market operator needs a temporary cloud-resilience leader to restore recovery evidence, clarify failover authority and complete a fourteen-month succession.
The mandate
A cloud control-plane failure prevented automated recovery of two electricity-market services and exposed a gap between application runbooks, provider assumptions and operational command. The resilience executive resigned during the subsequent review, leaving platform and business-continuity leaders with overlapping remits but no single authority for the recovery programme.
The interim must arrive within three weeks and serve for fourteen months, spending four days weekly in Sydney during the first quarter. A permanent search starts after the first full-scale exercise, with six weeks reserved for the selected successor to observe and then command rehearsals before taking accountability.
Handover is complete when every critical cloud-hosted service has a business-owned tolerance, a technically proven recovery path, two unannounced exercises meet those tolerances, concentration exceptions have dated treatments, and the successor has signed the next annual test plan. Documents without restored services and reproducible evidence will not close the mandate.
The leader may invoke technology continuity arrangements, require service isolation, reorder resilience engineering, withhold production releases that breach recovery gates and enforce contractual testing within approved spend. Architecture changes above AUD 8 million, movement of regulated data, permanent appointments and acceptance of tolerance breaches require executive or board approval; the interim cannot redefine market obligations.
This is not a wholesale cloud migration, data-centre closure or cyber-transformation assignment. Product feature roadmaps, electricity dispatch policy and enterprise application modernisation remain with their existing owners unless they directly prevent an agreed recovery test.
Why this seat is open
The outage disproved assurances that provider-native redundancy alone met the organisation's critical-service tolerances. An executive resignation then removed the person expected to reconcile supplier, engineering and operational accountability. The board needs a finite period of command, proof and transfer before entrusting the permanent leader with the next resilience cycle.
What you will own
- Reconcile business-impact tolerances with actual application dependencies, recovery mechanics, data-loss points and manual operating alternatives.
- Decide which services require redesign, independent fallback, degraded operation or documented risk acceptance based on tested recoverability.
- Establish a release gate that rejects material cloud changes lacking updated failure modes, rollback proof and named incident authority.
- Renegotiate supplier exercise obligations, evidence access, escalation paths and dependency disclosures within current commercial boundaries.
- Command two unannounced recovery scenarios that combine control-plane loss, regional impairment and constrained specialist availability.
- Publish board evidence linking each tolerance breach to an accountable treatment, funding decision, interim control and retest date.
- Induct the successor through a reversed-shadow exercise and transfer the service map, decision log, supplier posture and annual assurance calendar.
Candidate qualifications
- Held CIO-minus-one or enterprise director accountability for cloud operations and resilience in regulated critical infrastructure or essential services.
- Personally led recovery from a material cloud failure involving provider control-plane, identity, network or data-consistency complications.
- Converted business-impact tolerances into engineered recovery designs and witnessed tests across complex application dependency chains.
- Challenged hyperscale providers and major integrators on concentration, evidence and contractual exercise rights at executive level.
- Exercised command across technology, operations, communications, legal and regulatory stakeholders during high-consequence service scenarios.
- Left a permanent successor with a repeatable resilience governance cycle whose assurance did not depend on the interim's presence.
Non-negotiables
- Can start within three weeks, work four days weekly in Sydney initially and travel monthly to Melbourne or Brisbane.
- Will accept an exclusive executive commitment and join the out-of-hours critical-incident roster for the assignment.
- Has direct accountability for tested cloud recovery, not only policy, audit or migration programme management.
- Must disclose current hyperscaler, integrator and Australian critical-infrastructure board relationships before interview.
- 49 words maximum. What is your earliest Sydney start date, and which existing commitments must end before you hold this resilience seat?
- 49 words maximum. Describe a cloud recovery test that failed under your authority and the design decision the evidence forced.
- 49 words maximum. Which metric would cause you to withhold a production release despite commercial pressure to proceed?
This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.