Confidential mandate
etcd Quorum Recovery Architecture Director
Planned Hiring / New
etcd Quorum Recovery Architecture Director mandate in Istanbul, Türkiye · Digital Wholesale Marketplace
A wholesale marketplace needs a six-week recovery architecture after Kubernetes control-plane restores produced inconsistent etcd membership, encrypted state and revision histories across important production regions.
The mandate
Recovery tests for regional Kubernetes control planes produced etcd clusters with stale members, mismatched peer identity, unavailable encryption keys and revision histories that invalidated controllers’ assumptions. Snapshot completion had been treated as proof of restorable state, while compaction and external dependencies varied by cluster. The defined problem is to establish a reproducible quorum-loss and restore method with explicit consistency boundaries.
The named deliverable is an etcd Quorum Recovery Architecture, Experiment Corpus and Operator Runbook. It must cover membership and identity, snapshot integrity, encryption and key availability, revision behaviour, compaction, watches, leases, external control dependencies, topology, restore sequence, validation, rollback, evidence, tooling and bounded regional recovery patterns for the existing platform.
Three milestones govern six weeks: by the end of week one, accept the state and failure inventory; by week four, complete quorum-loss, corrupted-snapshot, key-loss and restored-revision experiments; and by week six, deliver the architecture, tested runbook, automation requirements and signed recovery opinion. Each milestone invoice follows acceptance of the related evidence.
The chief platform engineer and recovery board will accept the work only when snapshots prove integrity before use, membership and peer identity rebuild deterministically, encrypted objects remain recoverable under approved key-loss scenarios, controllers tolerate or explicitly handle restored revisions, client operators repeat two timed recoveries, and every external prerequisite or impossible condition appears in the runbook.
The client will provide cluster and etcd configurations, snapshots, encryption settings, key services, certificate material, controller inventories, backup automation, isolated cloud environments, incident histories and cleared operators. The consultant will not restore production, rewrite application controllers, replace the Kubernetes distribution, operate key infrastructure or certify business data; unavailable proprietary components remain stated recovery dependencies.
Why this is external work
The platform team wrote the backup automation and has repeatedly tested the successful path, while distribution support focuses on supported commands rather than the client’s external control dependencies. Recent inconsistent outcomes require an independent distributed-state diagnosis. A specialist can force quorum, identity, encryption and revision assumptions into one reproducible method before another regional event makes production the test environment.
What you will own
- Inventory members, peer identity, certificates, snapshots, encryption providers, keys, revisions, compaction, leases and dependent controllers.
- Define recoverable-state boundaries and prerequisites for member loss, quorum loss, region loss, corrupted snapshot and unavailable key service.
- Build experiments for stale membership, peer mismatch, partial snapshot, encryption-key absence, compaction and changed restore topology.
- Trace restored revisions, watches, leases and reconciliation through representative controllers and external automation.
- Specify validation for object counts, hashes, critical resources, encryption, control convergence, admission and marketplace readiness.
- Design operator sequence, independent evidence, stop conditions, rollback and escalation for each supported recovery pattern.
- Deliver the architecture, experiment scripts, timed runbook, automation backlog and residual impossibility register.
Candidate qualifications
- Recovered production etcd quorum and Kubernetes control state from concurrent member, region, snapshot or encryption failures under pressure.
- Diagnosed stale peer identity, restored revision, compaction, watch, lease or controller effects after snapshot recovery.
- Integrated key and certificate dependencies into control-plane recovery rather than assuming cryptographic services remain available.
- Built isolated failure experiments that distinguished a valid snapshot from an operationally recoverable cluster.
- Authored timed operator runbooks with validation, stop, rollback and impossible-state escalation.
- Delivered a recovery architecture that client operators repeated independently under observed fault conditions.
Non-negotiables
- Can complete two Istanbul recovery laboratories and the regional review inside the fixed six-week calendar.
- Will disclose relationships with Kubernetes distributions, cloud, backup, key-management and platform-recovery vendors.
- Brings hands-on etcd recovery evidence; Kubernetes application operations or generic backup work alone are insufficient.
- Accepts two client-run timed recoveries and explicit impossible conditions as final acceptance requirements.
- 49 words maximum. Describe an etcd snapshot that was valid but could not restore an operable Kubernetes control plane.
- 49 words maximum. Which controller behaviour would you test first after restoring to an earlier revision?
- 49 words maximum. How would you recover encrypted objects when quorum and the external key path fail together?
This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.