Confidential mandate
Regulated Chaos-Experiment Control Director
Planned Hiring / New
Regulated Chaos-Experiment Control Director mandate in Munich, Germany · Connected Medical Monitoring
A connected-medical platform needs a four-month control framework after inconsistent chaos experiments exposed patient-monitoring services across regions to failure combinations outside approved clinical and operational boundaries.
The mandate
Platform teams introduced chaos tooling to test recovery, but experiments differ in hypothesis, patient exposure, abort logic, evidence and quality approval. One exercise combined queue delay and regional database loss beyond its authorised boundary, causing remote-monitoring alerts to arrive outside their expected window. The defined problem is to make resilience experimentation safe, decision-bearing and repeatable in a clinically governed production estate.
The named deliverable is a Regulated Chaos-Experiment Control Framework and Validated Scenario Library. It must define experiment classes, clinical and operational risk assessment, prerequisites, blast-radius limits, patient and device exclusions, approvals, observability, abort and rollback, evidence retention, independent review, post-experiment learning, tooling boundaries and six executable reference scenarios.
Four milestones govern four months: by week two, accept the service and experiment inventory; by week six, deliver the risk taxonomy and control design; by week twelve, conduct six bounded experiments across non-production and approved production slices; and by week seventeen, submit the framework, scenario library, training package and governance recommendation. Invoices follow milestone acceptance.
The quality and technology officer with the acceptance board will approve the work only when every scenario states a falsifiable hypothesis and prohibited impact, preconditions are machine-checked where feasible, abort signals operate independently of the injected failure, patient and device exposure can be reconstructed, two client teams repeat selected experiments safely, and findings produce owned reliability decisions.
The client will provide service maps, clinical-risk classifications, incident records, existing experiment scripts, observability access, test tenants, device simulators, quality procedures and authorised production windows. The consultant will not conduct unapproved live experiments, change clinical algorithms, assess medical efficacy, buy tooling, operate incident response or waive quality controls; any untestable hypothesis remains outside the validated library.
Why this is external work
The SRE teams running experiments also define their success and have uneven familiarity with clinical quality obligations. Quality staff can constrain patient risk but lack deep failure-injection experience across distributed cloud dependencies. An independent specialist is required to create a common control language and demonstrate that useful experiments can remain bounded without becoming ceremonial tests.
What you will own
- Classify experiments by clinical consequence, customer exposure, reversibility, propagation potential, data integrity and required approval.
- Convert incident hypotheses into explicit steady state, injected condition, expected degradation, prohibited outcome, abort signal and learning decision.
- Define blast-radius controls across tenants, regions, devices, queues, databases, identity, networks, time and dependent suppliers.
- Establish independent observability and stop paths that remain available when the target service or primary telemetry channel fails.
- Run six reference scenarios spanning latency, dependency loss, corrupted response, stale state, capacity exhaustion and partial regional isolation.
- Produce evidence linking approvals, versions, exposure, observations, intervention, recovery, residual risk and accountable remediation.
- Deliver the control framework, scenario code, operator training, review templates and prioritised experiment roadmap.
Candidate qualifications
- Designed and governed production chaos engineering for regulated healthcare, finance, transport or comparably consequential services across multiple regions.
- Ran bounded live experiments with explicit customer or patient exclusions, independent abort and fully reconstructable exposure evidence.
- Converted severe past incidents into falsifiable hypotheses whose controlled tests drove architecture or operating decisions.
- Integrated reliability experimentation with quality, clinical, privacy, security and change-control functions.
- Stopped or redesigned an experiment when compound failure could exceed its declared blast radius.
- Delivered experiment libraries that client teams safely repeated without ongoing consultant supervision.
Non-negotiables
- Can complete both Munich experiment residencies and all four acceptance reviews inside the four-month term.
- Will disclose relationships with chaos, observability, cloud and connected-device platform vendors.
- Brings live regulated experimentation evidence; game days, tabletop facilitation or tooling demonstrations alone are insufficient.
- Accepts client repetition, independent abort and reconstructable exposure as final acceptance requirements.
- 49 words maximum. Describe a chaos experiment you stopped because its possible blast radius exceeded the approved hypothesis.
- 49 words maximum. Which abort channel would you keep independent when injecting failure into remote-monitoring alert delivery?
- 49 words maximum. How would you prove that no excluded patient or device entered an experiment cohort?
This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.