Confidential mandate

LLM Platform Reliability and FinOps Leader

Urgent / Replacement

LLM Platform Reliability and FinOps Leader mandate in Dublin, Ireland · Travel Marketplace Technology

A Dublin travel marketplace needs a twelve-month executive leader after an LLM routing outage and retry-driven spend surge, ending with workload-level reliability economics and an inducted permanent successor.

The mandate

The platform executive departed after a model-provider disruption caused the gateway to retry long-running customer-service tasks across progressively more expensive endpoints, exhausting regional queues while multiplying inference charges. Supplier-onboarding and disruption-support products degraded together, yet incident command could not identify which retry policy, context expansion or fallback choice drove cost because technical traces and financial allocation used different workload identities.

The interim must join in Dublin within fourteen days and hold executive platform authority for twelve months. The permanent search will start after ninety days of stable workload metering in month six, followed by a six-week overlap covering budget, incident and provider decisions; the seat ends at twelve months even if a product team has not migrated every legacy prompt.

Handover is achieved when each production workload has a service class, context and model policy, fail-safe degradation path, accountable budget and trace from request through retries to invoice; two consecutive quarters must meet agreed availability, latency and cost-to-serve bands without suppressing demand. The successor must command a provider-loss exercise, approve the next capacity plan and accept the exception and commitment registers before the committee closes the mandate.

The interim may suspend workloads, change routing and retry limits, define platform service levels, reassign on-call ownership, procure short-term capacity and commit up to €18 million within the approved recovery envelope. Committee approval is required for multi-year provider commitments, expenditure above that ceiling, changes to enterprise resilience appetite or withdrawal of a customer-facing capability; permanent executive hiring, product pricing and settlement of provider disputes remain reserved.

Redesign of travel products, customer-service policy, marketplace take rates and general cloud migration outside the LLM platform are out of scope. The leader will require product demand, prompt and revenue data to govern service classes, but does not own response content, model training on customer conversations, non-LLM infrastructure or commercial negotiation unrelated to platform capacity and reliability.

Why this seat is open

The outage exposed a platform optimised for provider choice but not for bounded failure across retries, queues, context growth and spend. Reliability engineers and Finance each hold partial evidence, while product teams can alter demand without accepting the resulting service and cost trade-offs. The board needs temporary joint technical-economic authority before appointing a permanent leader to a newly defined seat.

What you will own

  • Reconstruct the disruption from gateway decision, context size, timeout and retry through regional queue, provider response, user outcome and invoiced charge.
  • Classify workloads by consequence, latency, quality tolerance, data boundary and economic value, assigning explicit degradation and shutdown behaviour to each.
  • Decide routing, caching, batching, context, retry and fallback policies using measured end-to-end outcomes rather than model price or benchmark quality alone.
  • Build request-level cost attribution that reconciles provider, cloud, vector, observability and support charges to product, customer cohort and failed work.
  • Set capacity and commitment gates across reserved throughput, on-demand endpoints and secondary providers under demand, outage and model-change scenarios.
  • Reset change control so prompt, model, retrieval and gateway revisions carry load, reliability and cost evidence before production promotion.
  • Induct the permanent leader through a provider failure exercise, quarterly capacity decision and budget review, transferring authorities and unresolved dependencies.

Candidate qualifications

  • Held executive responsibility for a shared LLM or high-scale inference platform serving multiple products with explicit reliability and financial outcomes.
  • Recovered a model-routing or distributed-service incident involving retry amplification, queue saturation or failed fallback and can evidence the durable correction.
  • Joined request-level telemetry to provider and cloud invoices sufficiently to change routing, product demand or capacity commitments.
  • Designed workload service classes and graceful degradation where availability, latency, model quality, data controls and cost could not all be maximised.
  • Negotiated technical capacity and resilience choices across several model providers without allowing commercial incentives to determine architecture.
  • Completed a platform leadership transition after stabilising operations, redefining FinOps ownership and testing the successor through a live or simulated incident.

Non-negotiables

  • Available within fourteen days for an exclusive Dublin assignment and able to complete monthly Barcelona or Frankfurt plus quarterly provider travel.
  • Will shed or suspend lower-value demand before an uncontrolled retry pattern threatens critical workloads or the agreed spending envelope.
  • Has personally operated shared production inference at material scale; prototype deployments or cloud FinOps without LLM-platform accountability do not qualify.
  • Must disclose any investment, advisory role, rebate or other commercial interest involving a model, cloud, observability or optimisation provider under review.
  1. 49 words maximum. State your notice position, earliest Dublin start and constraints on the required regional and provider travel.
  2. 49 words maximum. Describe a retry or routing incident you recovered, including how you identified both the reliability and cost amplification path.
  3. 49 words maximum. How did request-level economics change one model, fallback, capacity or product-demand decision in a shared inference platform?

This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.