Confidential mandate

Multi-Cluster Workload Placement Control Director

Planned Hiring / New

Multi-Cluster Workload Placement Control Director mandate in Sydney, Australia · Digital Assessment Platforms

A digital assessment provider needs a five-month placement architecture after regional schedulers concentrated examination workloads during peak periods despite declared resilience, latency, residency and disruption constraints.

The mandate

Regional schedulers and deployment pipelines independently place examination services across clusters using affinity, disruption and capacity signals that conflict under peak load. A recent rehearsal concentrated candidate sessions and scoring dependencies inside one cloud failure domain despite documented resilience intent. The defined problem is to establish one explainable placement control model spanning latency, residency, state, failure and scheduled examination consequence.

The named deliverable is a Multi-Cluster Workload Placement Architecture and Conformance Harness. It must include service classes, hard and soft constraints, source authority, conflict precedence, data and session affinity, failure-domain budgets, disruption controls, capacity reservations, placement explanations, override governance, test scenarios, rollout sequence and an implementation specification for the existing orchestration stack.

Four milestones span five months: by week three, accept the service, cluster and constraint inventory; by week nine, deliver the placement and authority model; by week sixteen, complete replay and controlled-failure tests against two peak profiles; and by week twenty-two, submit the conformance harness, target architecture and rollout backlog. Each invoice follows acceptance of its milestone.

The platform officer and examination reliability board will accept the final work only when every critical placement decision can be explained from versioned inputs, hard residency and state rules cannot be overridden by stale capacity, simulated region and cluster losses preserve signed session thresholds, conflicting policies fail closed or escalate, and client engineers can add an unseen service to the harness without consultant support.

The client will provide cluster inventory, scheduler and deployment configuration, service topology, session and scoring flows, residency rules, disruption budgets, capacity history, peak profiles, incident evidence and test environments. The consultant will not operate live examinations, rewrite assessment software, buy orchestration tooling, negotiate cloud contracts or approve regulatory interpretations; missing application constraints remain recorded inputs.

Why this is external work

Platform and application teams authored different placement rules and each can show its component behaved as configured. Examination operations experience the combined outcome but lack authority to arbitrate technical precedence. External distributed-platform expertise is needed to convert business constraints into one testable control model before the next high-stakes assessment season.

What you will own

  • Catalogue workload state, latency, residency, dependency, disruption, scaling, capacity and failure requirements by examination journey.
  • Define authoritative sources and precedence for hard constraints, preferences, reservations, health, maintenance and emergency overrides.
  • Model placement across cluster, zone, region and provider failure domains with explicit correlated-dependency budgets.
  • Build conformance scenarios for peak admission, stale capacity, unavailable clusters, data-locality conflict, disruption and partial control loss.
  • Replay representative placement histories to identify concentration, thrashing, unschedulability and hidden policy overrides.
  • Specify decision explanations, audit evidence, approval, expiry and rollback for any manual or automated exception.
  • Deliver the architecture, harness, tested scenarios, rollout waves and unresolved constraint register at final acceptance.

Candidate qualifications

  • Designed production multi-cluster Kubernetes or comparable workload placement for stateful, regulated or peak-critical services across regions.
  • Resolved policy conflicts among residency, latency, affinity, capacity, disruption and correlated failure-domain requirements.
  • Built explainable scheduling or orchestration controls whose decisions could be replayed from versioned inputs.
  • Tested regional loss and peak admission together, exposing concentration that steady-state capacity views had missed.
  • Governed emergency placement overrides with expiry, evidence and safe restoration to policy control.
  • Delivered a conformance harness that client teams extended independently to new services.

Non-negotiables

  • Can complete both Australian peak-readiness residencies and all four milestone reviews within five months.
  • Will disclose relationships with cloud, Kubernetes, scheduler, deployment and assessment-platform providers.
  • Brings production multi-cluster placement evidence; cluster administration or capacity dashboards alone are insufficient.
  • Accepts client extension of the harness and explainable placement as final acceptance conditions.
  1. 49 words maximum. Describe a placement decision that satisfied capacity but violated a more important hidden constraint.
  2. 49 words maximum. Which combined peak and failure scenario would you replay before an examination season?
  3. 49 words maximum. How should a scheduler respond when residency evidence and available-capacity data disagree?

This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.