Confidential mandate
Dependency-Aware Error-Budget Architecture Director
Planned Hiring / New
Dependency-Aware Error-Budget Architecture Director mandate in São Paulo, Brazil · Instant Commerce Platforms
An instant-commerce network needs a four-month reliability model after teams met local service objectives while shared dependency exhaustion still broke complete ordering and fulfilment journeys.
The mandate
Service teams report healthy local objectives while customers still fail to place and receive orders when shared identity, catalogue, inventory, payment or dispatch dependencies degrade. Current error budgets neither allocate shared failure nor arbitrate which releases should stop, and financial impact is reconstructed after the fact. The defined problem is to create journey-level reliability budgets that remain actionable for autonomous teams.
The named deliverable is a Dependency-Aware Error-Budget Architecture and Operating Compact. It must include journey and dependency maps, indicator definitions, attribution rules, shared-budget mathematics, regional and peak segmentation, economic consequence, release and investment policies, dispute resolution, evidence quality, governance, dashboards and a tested quarterly operating cycle.
Four milestones govern four months: by week two, accept the journey, dependency and incident inventory; by week seven, deliver indicators and attribution proposals; by week twelve, replay six months of failures and releases through the model; and by week seventeen, submit the operating compact, measurement specification and adoption backlog. Each invoice follows acceptance of its named output.
The operating and technology officer with the acceptance council will approve the work only when journey indicators reproduce sampled customer outcomes, shared dependency failures allocate without double counting, teams can explain release decisions from versioned evidence, economic thresholds distinguish region and peak consequence, and client leaders complete one budget review and resolve a disputed attribution without consultant intervention.
The client will provide service maps, traces, incident records, orders, fulfilment outcomes, release histories, cost and margin data, regional calendars, current objectives and responsible engineers. The consultant will not operate production, set commercial service promises, approve releases, purchase observability tools or reorganise teams; unavailable customer evidence remains a stated confidence limitation.
Why this is external work
Each service owner benefits from a local measure that excludes dependency and customer failure, while finance lacks the technical basis to attribute shared impact. Previous attempts became negotiations over percentages rather than evidence. An external SRE economics practitioner is needed to create a coherent model and facilitate its first use without inheriting any team’s existing performance position.
What you will own
- Map ordering, payment, allocation, picking, dispatch and delivery journeys to shared and regional service dependencies.
- Define customer-valid indicators for completion, correctness, timeliness, duplication, compensation and recovery rather than endpoint availability alone.
- Design attribution for shared failures, correlated incidents, retries, partial journeys and overlapping budget consumption without double counting.
- Segment budgets by geography, peak, customer promise and economic consequence while preserving one enterprise decision language.
- Replay incidents and releases to test whether proposed thresholds would have changed action at the time evidence existed.
- Specify release stop, recovery priority, reliability investment, exception and dispute rules tied to measurable budget state.
- Deliver the architecture, operating compact, measurement contracts, replay corpus and first client-led review evidence.
Candidate qualifications
- Designed error budgets or SLO systems across multi-service customer journeys with shared dependencies.
- Connected technical reliability to order, fulfilment, revenue or customer consequence without reducing decisions to uptime.
- Built failure-attribution methods that avoided double counting and survived challenge from autonomous service owners.
- Replayed historical incidents and releases to validate whether proposed policy would have produced better decisions.
- Facilitated contested reliability investment or release choices using explicit economic and evidence thresholds.
- Delivered an operating cycle that client leaders ran independently after the consulting engagement.
Non-negotiables
- Can complete two Brazilian operations residencies and all four acceptance reviews within four months.
- Will disclose observability, commerce-platform, cloud and SRE-service relationships before receiving performance data.
- Brings journey-level reliability economics; isolated service SLO implementation alone is insufficient.
- Accepts client-led attribution and release-decision rehearsal as final acceptance conditions.
- 49 words maximum. Describe a journey failure where every contributing service still met its local objective.
- 49 words maximum. How would you allocate one shared dependency incident without charging the same customer harm twice?
- 49 words maximum. Which historical release decision would you replay first to validate an enterprise error budget?
This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.