Confidential mandate

RDMA Fabric Congestion Evidence Director

Planned Hiring / New

RDMA Fabric Congestion Evidence Director mandate in Doha, Qatar · Subsurface Energy Modelling

An energy-modelling centre needs a ten-week independent diagnosis after distributed simulations showed severe operational tail stalls despite acceptable average RDMA bandwidth and deceptively low host utilisation.

The mandate

Large reservoir simulations exhibit episodic collective-operation stalls while average fabric bandwidth, node use and packet-loss dashboards remain inside target. Teams disagree whether incast, path imbalance, lossless Ethernet controls, host pacing, storage bursts or application synchronisation is causal. The defined problem is to produce an evidence-based congestion diagnosis and bounded remediation for the representative production workload.

The named deliverable is an RDMA Fabric Congestion Evidence Dossier and Remediation Design. It must include workload communication signatures, topology and path map, counter and timestamp evidence, queue and congestion-control behaviour, host and application interaction, controlled experiments, identified causal chains, configuration changes, observability requirements, rollback, and a scaled validation plan.

Three milestones govern ten weeks: by week two, accept the topology, workload and measurement baseline; by week six, complete controlled reproduction across traffic and fault conditions; and by week ten, deliver the diagnosis, tested remediation, operating thresholds and final evidence dossier. Each milestone invoice follows approval of its named outputs.

The computing officer and HPC acceptance panel will sign only when representative tail stalls reproduce on demand, evidence aligns fabric, host and application time, alternative hypotheses are tested rather than asserted, proposed changes improve signed tail and completion thresholds without unsafe loss or deadlock, and client engineers repeat both the diagnosis and rollback using the delivered procedure.

The client will provide fabric topology, device and host configuration, switch and NIC counters, packet and congestion telemetry, representative simulations, MPI traces, storage activity, test partitions, vendor support and performance engineers. The consultant will not rewrite scientific solvers, procure network equipment, operate production jobs, change facility power or certify vendor products; unobservable mechanisms remain limitations.

Why this is external work

Network, host, storage and simulation teams each possess a plausible explanation aligned with the layer they own. Vendor specialists can tune their components but are not neutral about architectural or configuration conclusions. External RDMA performance expertise is required to synchronise evidence across layers and reject attractive hypotheses that do not reproduce the long-tail failure.

What you will own

  • Characterise communication, message size, collective timing, burst, synchronisation and storage interaction for representative simulation phases.
  • Map physical and logical paths, oversubscription, adaptive routing, priority, buffer and failure domains across the fabric.
  • Align switch, NIC, host, MPI, accelerator and storage evidence on a defensible time base during tail stalls.
  • Design experiments for incast, path imbalance, congestion notification, pause propagation, receiver pressure, host pacing and background traffic.
  • Isolate causal chains through controlled configuration changes, counterfactual workloads and rollback rather than correlation alone.
  • Define operating thresholds and telemetry for queue growth, pause, marking, retransmission, path skew and completion-time risk.
  • Deliver the diagnosis, configuration design, test corpus, rollback procedure and scaled validation plan at final acceptance.

Candidate qualifications

  • Diagnosed production RDMA or comparable low-latency fabric congestion in large HPC, AI or storage clusters.
  • Correlated switch, NIC, host, MPI and workload evidence to explain severe tails hidden by average bandwidth.
  • Tested incast, congestion control, lossless behaviour, path imbalance and receiver pressure through controlled experiments.
  • Distinguished network causality from application synchronisation and storage bursts using repeatable counterfactuals.
  • Changed fabric or host configuration with explicit deadlock, loss, fairness, rollback and workload-regression controls.
  • Delivered a performance diagnosis that client engineers reproduced independently at representative scale.

Non-negotiables

  • Can work on site in the Doha compute laboratory and complete the supplier review within ten weeks.
  • Will disclose ties to switch, NIC, server, accelerator, storage and HPC software providers.
  • Brings hands-on RDMA congestion evidence; general network design or benchmark reporting alone is insufficient.
  • Accepts repeatable tail reproduction, alternative-hypothesis testing and client-run rollback as final acceptance conditions.
  1. 49 words maximum. Describe an RDMA tail stall that average bandwidth and host utilisation failed to reveal.
  2. 49 words maximum. Which aligned counters would you collect first to distinguish incast from application synchronisation?
  3. 49 words maximum. How would you test a lossless-fabric tuning change for new deadlock or fairness risk?

This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.