Confidential mandate

Product-Experiment Evidence Director

Planned Hiring / New

Product-Experiment Evidence Director mandate in Helsinki, Finland · Mobile Gaming Technology

A mobile games publisher needs a four-month repair of experimentation evidence after identity gaps, overlapping releases and selective metrics made product investment decisions impossible to reproduce.

The mandate

Studios run hundreds of live experiments, yet product leaders cannot reproduce why several economy and progression changes were scaled. Device identity resets alter assignment, concurrent events contaminate cohorts, late revenue is attributed selectively and teams promote whichever engagement horizon supports their preferred release. The defined problem is to restore credible causal evidence quickly enough for live operations without forcing creative teams through a central statistics gate for every low-risk iteration.

The deliverables are an experiment-failure anatomy, exposure and identity contract, metric-and-guardrail catalogue, interference controls, review protocol, investment evidence standard and remediation roadmap. The design must address cross-device players, geographic eligibility, seasonality, network effects, whale concentration, delayed monetisation, novelty, repeated testing and consent-driven data gaps. It must distinguish directional discovery from decisions that alter a title’s economy or portfolio funding.

Three milestones organise four months: week four accepts the forensic review of twelve disputed experiments; week ten approves the target assignment, exposure, metric and decision rules after replay tests; and week seventeen delivers calibrated studio adoption, an independently reproduced scale decision, tooling requirements and the investment-council assurance paper. Fees are billed against those three milestones once their evidence is signed.

Acceptance requires analysts outside the originating studio to recreate assignment, exclusions, exposure, primary outcomes and uncertainty for the selected cases from governed data. Studio heads must classify experiments consistently, and the investment council must be able to distinguish learning from claimed value. Guardrails must expose harm to new, returning, high-spend and vulnerable cohorts rather than letting aggregate engagement conceal redistribution.

The client provides raw event schemas, identity stitching logic, assignment services, release calendars, experiment notebooks, economy changes, revenue maturation data, privacy constraints and access to studio decision-makers. The consultant does not choose game content, operate live events, write production analytics code, approve player targeting, make privacy determinations or decide portfolio funding. Client teams implement controls and retain all release authority.

Why this is external work

Each studio has evolved its own valid-looking analytical conventions, and central data teams helped build the services now under question. Reconciliation has stalled because disagreement is treated as statistical preference rather than product-governance risk. Independent expertise can replay contested cases, separate legitimate contextual variation from selective method and define proportionate evidence that creative leaders will use without turning experimentation into theatre.

What you will own

  • Forensically replay disputed experiments from assignment and exposure through exclusions, outcomes, maturity windows and scale decisions.
  • Define identity and exposure contracts for cross-device play, reinstall behaviour, account linking, consent gaps and shared households.
  • Establish primary measures and cohort guardrails covering progression, economy health, retention, monetisation, fairness and player harm.
  • Design controls for concurrent releases, seasonality, interference, novelty, repeated looks, delayed value and concentrated spend.
  • Create evidence tiers separating exploratory learning, reversible live tuning, economy change and portfolio-capital decisions.
  • Calibrate studio and central analysts through blinded case review, explicit disagreement and independently reproduced conclusions.
  • Deliver decision records, metric ownership, tooling requirements, remediation priorities and a sustainable assurance sampling rhythm.

Candidate qualifications

  • Has governed high-volume experimentation for mobile games, consumer platforms or another product with networked and repeated behaviour.
  • Can diagnose assignment, exposure, identity, interference, maturation and multiple-testing failures from underlying events and code logic.
  • Understands live-game economies, cohort concentration and why short engagement lifts can damage retention, trust or long-term monetisation.
  • Has created proportionate experiment evidence tiers that preserve team speed while strengthening consequential product and capital decisions.
  • Can facilitate contested causal reviews among statisticians, studio leaders, economy designers, finance and privacy specialists.
  • Leaves reproducible practices and accountable metric stewardship rather than a consultant-owned scoring model or generic analytics maturity map.

Non-negotiables

  • Can attend the Helsinki design cadence, both studio residencies and the cross-title calibration council within four months.
  • Will disclose relationships with publishers, studios, experimentation vendors, analytics platforms and major gaming investors.
  • Brings forensic experiment-governance evidence; dashboard design or general data strategy alone is insufficient.
  • Will not validate a scale decision whose assignment, exposure or outcome population cannot be independently reconstructed.
  1. 49 words maximum. Describe an experiment whose apparent win disappeared when you reconstructed actual feature exposure.
  2. 49 words maximum. Which guardrail would reveal that an economy change benefits averages by harming new players?
  3. 49 words maximum. How would you classify a live-game test that is reversible technically but not economically?

This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.