Confidential mandate
Chief Data Officer — Machine-Learning Infrastructure Stack
Planned Hiring / New
CDO - Data mandate in Zurich, Switzerland · Artificial Intelligence
Govern the data economics of a Zurich machine-learning stack, making corpus lineage, storage, movement and reuse visible before the next model-scale increase.
The mandate
A machine-learning infrastructure stack is storing and moving data at a rate that outpaces model-cost improvements. Multiple training corpora, derived features and evaluation copies have uncertain reuse, lineage and retention. Researchers need speed, but the organisation cannot explain which data assets drive model gain or recurring infrastructure cost. The board has created a planned CDO seat to establish data authority before the next scale step.
Approximately 400 employees and material partners span data engineering, infrastructure, research, security, privacy, product and operations from Zurich. The onsite CDO owns data strategy, governance, data products, quality, lineage, responsible use and data capability, reporting to the Group Chief Executive or designated executive committee sponsor. Research and model-release decisions remain with their authorised leaders.
The first inventory will identify logical assets rather than count files. Source, rights, transformations, versions, model use, evaluations, storage tier and accountable owner must connect. The CDO will expose duplicate or orphaned collections and distinguish a defensible independent snapshot from uncontrolled replication.
Lineage needs to work at training scale. Teams must be able to establish which sources and filters produced a corpus, which code transformed it and which model consumed it. The design should avoid impossible row-level overhead where aggregate provenance suffices, while preserving granular evidence for restricted or contested material.
Data rights will be operationalised. Licence, consent, residency, purpose, deletion and model-use conditions need machine-readable representation where practical and clear human decision routes elsewhere. A data asset cannot be approved simply because it is technically accessible. Downstream derivatives must retain relevant restrictions.
Storage economics will reflect access pattern and recovery need. Hot, warm, archive and deletion choices should consider retraining, reproducibility, regulatory hold and egress. The CDO will require an owner and expected use for premium storage. Moving cost between cloud accounts or teams does not count as saving.
Data movement is often hidden model cost. Cross-region transfer, repeated preprocessing and accelerator staging consume time and capacity. Architecture will favour locality, reusable validated products and planned caching without creating uncontrolled copies. Engineering measures should show end-to-end experiment cost, not compute alone.
Quality will be defined by use. Coverage, duplication, contamination, labelling, drift and sensitive-content risk can affect training and evaluation differently. Product owners will publish thresholds and known limitations. A central quality score will not replace a researcher's understanding of the failure being tested.
Evaluation data requires stronger separation. Leakage from training can produce compelling but invalid performance. Access, version and contamination checks will be independently visible. The CDO will ensure convenience copies do not blur the boundary between development and final assessment.
Deletion and correction need evidence. Removing source material from active stores may not remove it from derivatives, caches or future training queues. The organisation will define technical reach, proportional action and legal judgement for each request. Claims about model unlearning will not exceed what has been verified.
Data products should reduce repeated preparation. Curated corpora, feature sets, taxonomies and evaluation suites need service owners, documented applicability and adoption measures. The CDO will retire products whose maintenance cost exceeds demonstrated reuse, even if their catalogue presence appears strategically impressive.
Model-cost governance will link data choices to gain. Larger or cleaner corpora, additional languages and synthetic augmentation require hypotheses and comparison. Finance, research and data leaders will review marginal performance, risk and lifetime cost. Volume accumulated without an expected learning outcome will not be treated as an asset.
The data office must preserve research velocity through pre-approved paths, rapid rights advice and observable platforms. Governance will be judged by fewer late surprises and faster defensible reuse. Site-based data stewards and engineers will own execution; a policy team alone cannot control this estate.
What you will own
- Enterprise training and evaluation data strategy.
- Corpus inventory, lineage and accountable ownership.
- Rights, residency, retention and deletion controls.
- Storage, transfer and preparation economics.
- Use-specific data quality and contamination assurance.
- Reusable data products and lifecycle decisions.
- Data contribution to model-cost governance.
- Data organisation, stewardship and succession.
The first 12 months
Within 45 days, map priority training and evaluation assets, quantify storage and movement exposure and quarantine any collection lacking defensible rights or ownership. Establish interim approval for new large-scale ingestion.
By month six, deploy lineage for critical corpora, separate final evaluation assets and introduce tiering, retention and data-product ownership. Put marginal data value into model-investment reviews.
At twelve months, reduce avoidable storage and transfer cost by 35%, bring 95% of active training assets under verified lineage and retire 40% of orphaned copies. Every final evaluation suite must pass contamination controls, and deletion requests must have traceable disposition without unsupported claims of model removal.
What the sponsor will examine
- Logical assets governed beyond physical file counts.
- Rights surviving through transformations and derivatives.
- Data movement visible in experiment economics.
- Training and final evaluation remaining separated.
- Reuse measured before products are maintained indefinitely.
- Governance accelerating defensible research work.
The person
You bring 18–22 years in data leadership across machine learning, scientific computing, cloud platforms or regulated information estates. Your record includes CDO-scale accountability, large training corpora, lineage, rights, storage economics, evaluation integrity and technical teams operating internationally.
Candidates must discuss a costly data estate they reduced without losing reproducibility, and a corpus they withheld because provenance or permissible use was unclear. This permanent role is onsite in Zurich with close daily engagement across research and infrastructure.
Compensation and terms
Base compensation is CHF 450,000–620,000 plus annual incentive and LTI linked to lineage coverage, responsible use, model-cost improvement, research enablement and data capability. The permanent onsite Zurich CDO reports to the Group Chief Executive or designated executive committee sponsor. This is a planned new appointment.
Confidentiality
The organisation, corpora, models, rights, infrastructure costs, evaluation assets and control findings remain confidential. Further access follows credential assessment, conflicts and signed confidentiality. Candidates must not contact research laboratories, data suppliers, cloud providers or employees to identify the client.
More seats like this one
This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.