Confidential mandate
Synthetic Financial Data Fidelity Director
Planned Hiring / New
Synthetic Financial Data Fidelity Director mandate in Paris, France · Wholesale Banking Data Engineering
A Paris wholesale bank commissions a fourteen-week fidelity assessment to determine which synthetic transaction datasets preserve rare financial behaviour, protect confidentiality and can safely support three defined analytical uses.
The mandate
The bank needs a defensible answer for six synthetic transaction datasets proposed for financial-crime model development, market-stress analysis and non-production software testing. Current vendor reports show realistic marginal distributions, yet internal users have found broken event order, weakened tail dependence and missing rare typologies; privacy testing, meanwhile, does not distinguish memorised customer fragments from records that are merely plausible.
The commissioned artefact is a Synthetic Financial Data Fidelity Dossier and Use-Constraint Matrix. It will pair each dataset and purpose with temporal, relational, tail, privacy and downstream-task evidence; reproduce generator conditions; identify material blind spots; and state whether use is permitted, conditionally permitted or prohibited without implying that one aggregate similarity score certifies every analytical purpose.
Delivery begins 11 January 2027. Milestone one, due 22 January, is the reconciled dataset inventory and signed fitness questions; milestone two on 19 February is the metric harness and baseline challenge; milestone three on 19 March is the downstream utility, rare-event and disclosure-risk result set; milestone four on 16 April is the approved dossier, executable test pack and control handover.
Acceptance requires Model Risk to rerun the metric suite from retained code within agreed numerical tolerances, three domain owners to recover the signed sequence and tail behaviours from blinded samples, and Privacy Engineering to reproduce the membership and nearest-neighbour attack results. The joint sponsors must approve every use classification, caveat and retest trigger; generation-platform selection and production model approval remain separate decisions.
The bank will provide protected source samples inside its Paris analytics enclave, schema and relationship definitions, generator versions and parameters, downstream model code, rare-event typologies, existing privacy thresholds and named data stewards. A validation lead, financial-crime specialist and markets quant will attend weekly; the consultant may inspect Frankfurt validation evidence but may not remove row-level records, retrain production decision models or redesign the enterprise data platform.
Why this is external work
The generator vendors assess realism against their own methods, while internal teams advocate for the datasets that would remove their access bottlenecks. The bank lacks a specialist who can test temporal finance behaviour, statistical disclosure risk and downstream model utility without collapsing them into one score. A neutral engagement gives Model Risk a purpose-specific evidence record before synthetic data becomes embedded in regulated workflows.
What you will own
- Fix the six-dataset scope and translate each proposed use into observable fidelity, privacy, lineage and reproducibility questions before inspecting vendor scores.
- Reconstruct payment, order, trade and account-event sequences to expose impossible chronology, severed entity relationships and generator-induced state transitions.
- Measure tails, conditional dependence, regime behaviour and rare typology coverage using domain-signed tests that cannot be satisfied by marginal resemblance alone.
- Execute membership, attribute-inference, nearest-neighbour and canary attacks, separating disclosure evidence from unsupported claims that synthetic records are inherently anonymous.
- Compare downstream model ranking, calibration, stability and error concentration when trained or tested on synthetic, protected-real and deliberately degraded reference sets.
- Classify each dataset-use pairing as allowed, bounded or barred, attaching minimum sample controls, human review, retesting frequency and explicit prohibited extrapolations.
- Transfer the runnable harness, generator fingerprints, result lineage and exception register so validation can challenge a refreshed dataset without retaining the consultancy.
Candidate qualifications
- Directed synthetic-data validation for banking, payments or capital-markets records where sequence, entity relationship and extreme-event behaviour mattered operationally.
- Built fidelity assessments that combined statistical structure with domain event logic and downstream-task performance rather than relying on generic distance measures.
- Performed empirical disclosure-risk attacks on generated tabular or sequential data and communicated the remaining privacy uncertainty to accountable risk owners.
- Diagnosed a synthetic dataset that preserved common patterns but damaged tail dependence, rare typologies or temporal causality, leading to a documented use restriction.
- Delivered reproducible model-validation tooling inside a bank-controlled environment with traceable inputs, versioned generators and independent reperformance.
- Defended a contested permitted-use opinion before data, privacy and model-risk leaders whose incentives favoured faster access to production-like information.
Non-negotiables
- The proposed director must perform the core metric design and final use opinions, attend all Paris reviews and travel to the Frankfurt challenge workshop.
- No protected row-level data, generator artefacts or attack outputs may leave the bank environment or be submitted to an external analytical service.
- Any commercial relationship with a synthetic-data vendor, model platform or assurance supplier in scope must be declared before mobilisation.
- Will issue a prohibited-use conclusion when evidence fails, even if the affected dataset has already been purchased or promised to an internal programme.
- 49 words maximum. Describe a synthetic financial dataset whose common distributions looked credible but whose sequences, tails or entity relationships failed your validation.
- 49 words maximum. Which three tests would you prioritise to separate downstream utility from privacy leakage during the first four weeks, and why?
- 49 words maximum. Confirm fourteen-week capacity, Paris and Frankfurt attendance, and any relationship with a generator or platform provider relevant to this scope.
This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.