Confidential mandate
CXL Memory-Pooling Economics and Architecture Expert
Planned Hiring / New
CXL Memory-Pooling Economics and Architecture Expert mandate in Taipei, Taiwan · Advanced Server Systems
An advanced server manufacturer needs a fourteen-week independent recommendation on whether CXL memory pooling improves accelerator economics once latency, failure domains, software readiness and operations are counted.
The mandate
The manufacturer must decide whether pooled CXL memory belongs in its next accelerator-server platform or remains a limited customer option. Architecture models show higher memory utilisation, but prototype results vary by access locality, fabric contention and software behaviour; reliability teams question enlarged fault domains; and product finance values reclaimed capacity without including switch, firmware, qualification and support complexity. The defined problem is to test the complete proposition against real workloads.
The named deliverable is a CXL Memory-Pooling Architecture, Workload and Economics Recommendation. It must contain workload selection, topology alternatives, latency and bandwidth evidence, tiering behaviour, failure and containment results, firmware and operating-system dependencies, telemetry requirements, service model, bill of materials, lifecycle cost, ecosystem risks and a clear productise, constrain or defer conclusion.
Four milestones govern fourteen weeks: by week two, approve the assumptions and workload matrix; by week six, complete baseline and pooled-memory performance experiments; by week ten, finish failure, recovery, manageability and compatibility tests; and by week fourteen, submit the product recommendation, reference topology and qualification backlog. The matching milestone invoice follows acceptance of each output.
Acceptance rests with the product officer and architecture review board. They will sign only when test configurations and firmware are reproducible, benefits are segmented by workload rather than averaged, tail latency and fabric contention meet stated envelopes, poison and component-failure behaviour is demonstrated, the economics include stranded and support costs, and client engineers can reproduce the principal results independently.
The client will provide prototype hosts, CXL devices and switches, firmware, operating systems, accelerator platforms, benchmark and customer-derived workloads, power evidence, component costs, reliability engineers and laboratory access. The consultant will not design production silicon, negotiate suppliers, certify standards compliance, operate customer workloads or commit the product roadmap; unavailable ecosystem features remain explicit dependencies.
Why this is external work
Architecture teams sponsoring the platform and product teams seeking differentiation both benefit from a favourable answer. Component suppliers can demonstrate their part but cannot independently evaluate the enlarged system and support boundary. External expertise is required to unite memory behaviour, fabric failure, software maturity and product economics before platform decisions become locked into the next hardware cycle.
What you will own
- Select representative inference, training, HPC and data workloads by capacity pressure, locality, access pattern, sensitivity and operational value.
- Compare direct-attached, expanded, pooled and tiered topologies using throughput, median and tail latency, utilisation, power and workload completion.
- Stress fabric contention, hot pages, migration, oversubscription, topology asymmetry and noisy-neighbour behaviour under mixed demand.
- Test device removal, switch loss, link degradation, poison propagation, firmware restart and partial capacity recovery with attributable evidence.
- Define required discovery, allocation, telemetry, isolation, maintenance and incident controls across firmware, operating system and orchestration.
- Model fully burdened economics including memory reclaim, switching, cabling, power, qualification, spares, support, stranded capacity and refresh.
- Deliver the reference topology, results corpus, decision paper, product guardrails and qualification backlog at final acceptance.
Candidate qualifications
- Architected or evaluated CXL, memory pooling, tiering or composable infrastructure on physical multi-host systems.
- Measured workload locality, tail latency, bandwidth contention and capacity benefit without relying on synthetic averages alone.
- Tested memory-device, link, switch, poison or firmware failure and traced effects through operating systems and applications.
- Assessed ecosystem maturity across silicon, firmware, operating systems, orchestration, telemetry and field-support boundaries.
- Built hardware-platform economics incorporating qualification, service, power, stranded components and product lifecycle.
- Produced an independent product recommendation that survived review by architecture, reliability, finance and customer-engineering leaders.
Non-negotiables
- Can work from the Taipei integration laboratory and attend the stated supplier reviews during fourteen weeks.
- Will disclose relationships with memory, processor, switch, firmware, server and accelerator suppliers before testing.
- Brings hands-on CXL or comparable coherent-memory fabric evidence; conventional server sizing alone is insufficient.
- Accepts reproducible physical tests, failure results and client reruns as conditions for final acceptance.
- 49 words maximum. Describe a workload whose apparent pooled-memory capacity gain disappeared once locality or tail latency was measured.
- 49 words maximum. Which CXL failure test would you run before trusting a shared pool across accelerator hosts?
- 49 words maximum. How would you separate memory reclaimed economically from capacity made unusable by fabric or support overhead?
This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.