Confidential mandate

eBPF Observability Cost-and-Coverage Director

Planned Hiring / New

eBPF Observability Cost-and-Coverage Director mandate in Paris, France · Real-Time Advertising Exchange

A real-time advertising exchange needs a six-month observability design after telemetry costs rose sharply while consequential kernel, network and auction-path failures still escaped timely engineering detection.

The mandate

Telemetry volume and vendor charges have more than tripled, yet several lost-auction incidents were diagnosed only after revenue analysis because kernel queuing, packet drops and short-lived service behaviour were absent from application traces. Teams respond by adding high-cardinality labels, increasing cost without an explicit evidence model. The defined problem is to determine which eBPF and conventional signals justify their collection across the sub-100-millisecond auction path.

The named deliverable is an eBPF Observability Cost-and-Coverage Architecture and Migration Plan. It must contain an incident-to-evidence map, instrumentation design, cardinality budget, sampling and aggregation rules, privacy boundary, overhead envelope, storage tiers, vendor-neutral data contracts, operating ownership, experiment results and a phased retirement plan for redundant telemetry.

Four milestones govern six months: by week three, accept the incident corpus and current cost baseline; by week ten, deliver the signal and coverage model; by week eighteen, complete production-shadow experiments for kernel, network and runtime telemetry; and by week twenty-six, submit the target architecture, migration backlog and investment decision. Each milestone is invoiced only after its evidence review.

Acceptance belongs to the chief platform officer and finance partner. They will approve the final artefact only when ninety-five per cent of telemetry spend reconciles to a signal family, the top twelve failure hypotheses map to tested evidence, overhead stays within signed host and latency budgets, sensitive payloads are excluded by demonstrable controls, and client teams can apply the cardinality policy without consultant arbitration.

The client will provide telemetry invoices, schemas, retention rules, incident timelines, auction traces, representative hosts, kernel versions, privacy classifications and engineers for safe experiments. The engagement excludes buying an observability platform, operating production monitoring, rewriting auction logic, incident response outside agreed tests and renegotiating vendor contracts; inaccessible kernel environments will remain declared coverage gaps.

Why this is external work

Internal teams authored both the current instrumentation and the cost allocations now under challenge. Vendor specialists can optimise their own products but cannot neutrally recommend removing billable data or preserving portable evidence. An external practitioner is needed to join kernel mechanics, auction consequence and telemetry economics without turning the review into a generic tooling comparison.

What you will own

  • Reconcile ingestion, indexing, query, retention, egress and engineering costs to signal families, teams, environments and incident use.
  • Map auction failure hypotheses across kernel scheduling, sockets, packet paths, service runtime, queues and dependencies to minimum sufficient evidence.
  • Design eBPF probes and aggregation boundaries with verified overhead, kernel compatibility, payload exclusion and failure behaviour.
  • Set cardinality, sampling, exemplars, retention and escalation rules according to diagnostic value, rarity and commercial consequence.
  • Run shadow experiments comparing proposed signals with existing logs, metrics and traces during representative load and injected faults.
  • Define vendor-neutral event contracts, ownership, quality checks and graceful degradation when collectors or storage tiers fail.
  • Deliver the accepted architecture, economic model, migration waves, removal register and ongoing evidence-coverage scorecard.

Candidate qualifications

  • Designed eBPF-based production observability for high-throughput, latency-sensitive Linux infrastructure rather than security-only monitoring.
  • Connected kernel, network and runtime evidence to application or commercial failures that conventional traces had missed.
  • Reduced high-cardinality telemetry cost while preserving or improving incident-diagnostic coverage through explicit evidence design.
  • Measured probe, collection and export overhead under realistic host and workload conditions across heterogeneous kernels.
  • Established privacy and payload boundaries for low-level telemetry using tests rather than policy statements alone.
  • Delivered a vendor-portable observability architecture and retirement plan accepted jointly by engineering and finance leaders.

Non-negotiables

  • Can complete both Paris engineering residencies and all milestone reviews within six months.
  • Will disclose observability-platform, cloud, eBPF product and reseller relationships before receiving cost data.
  • Brings hands-on kernel instrumentation and telemetry-economics evidence; dashboard strategy alone is insufficient.
  • Accepts measured diagnostic coverage, overhead and client-operable cardinality control as final acceptance conditions.
  1. 49 words maximum. Which eBPF signal exposed a production failure that application tracing could not, and why?
  2. 49 words maximum. How would you value a high-cardinality field before deciding whether to retain, aggregate or remove it?
  3. 49 words maximum. Name the first overhead and privacy tests you would run against a proposed kernel probe.

This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.