Confidential mandate
AI Code-Migration Benchmark Director
Planned Hiring / New
AI Code-Migration Benchmark Director mandate in Warsaw, Poland · Life Insurance Core Systems
A Warsaw life insurer commissions a sixteen-week benchmark of three AI migration tools against representative legacy modules, requiring reproducible semantic, operational and maintainability evidence before a platform award.
The mandate
The insurer must decide whether any of three AI-assisted migration products can translate a defined portfolio of COBOL and PL/I policy, billing and actuarial modules into its target Java platform without moving hidden defects into a less intelligible form. Vendor demonstrations use small self-contained programs; the real estate contains packed decimals, copybook variants, batch restart logic, date conventions, embedded controls and undocumented behaviour that existing tests do not fully describe.
The engagement will deliver an AI Code-Migration Benchmark and Acceptance Harness. It will include a stratified source portfolio, behavioural oracle, data and interface fixtures, vendor-isolated execution protocol, semantic-difference register, performance and recovery tests, maintainability and security assessment, traceable scoring model and a procurement recommendation that separates automation capability from human remediation effort.
The sixteen-week clock starts 11 January 2027. Milestone one, the signed module population and behaviour inventory, is due 22 January; milestone two, the frozen harness, scoring rules and vendor environments, is due 19 February; milestone three, the blinded migrations and independently rerun results, is due 26 March; milestone four, the final benchmark, exception analysis and selection paper, is due 30 April.
Acceptance requires legacy owners to reproduce source behaviour from retained fixtures, target architects to build and run each submitted result in a clean environment, and actuarial control owners to reconcile precision-sensitive outputs under normal, boundary and restart cases. Cyber Security must reproduce the critical findings, Procurement must trace every weighted score to evidence, and both sponsors must sign unresolved behaviour, human-effort assumptions and conditions attached to the recommendation.
The client will provide source and build dependencies, compiler settings, job-control scripts, copybooks, interface contracts, production-derived masked fixtures, historical defects, batch logs, target standards and isolated vendor workspaces. Named legacy and actuarial experts will answer behaviour questions through a controlled channel; the consultancy will not execute the full migration, negotiate commercial terms, retire the mainframe or certify modules outside the benchmark sample.
Why this is external work
Legacy teams fear that an automation decision discounts tacit knowledge, while modernisation leaders and vendors are rewarded for demonstrating rapid conversion. Internal quality suites also encode today’s behaviour incompletely and cannot serve as an unquestioned oracle. Independent benchmark direction provides political neutrality, method control and a visible accounting of the expert labour hidden behind generated code.
What you will own
- Select a representative module portfolio across business criticality, language complexity, data shape, batch interaction, restart behaviour, embedded controls and defect history.
- Reconstruct behavioural oracles from code, tests, masked records, operational logs and expert decisions, marking uncertainty where no authoritative rule exists.
- Freeze vendor environments, prompts, tools, human-intervention limits and elapsed-time measures so each product encounters an equivalent migration problem.
- Execute semantic comparisons covering packed-decimal precision, date boundaries, error handling, sequencing, side effects, interfaces and interrupted batch recovery.
- Measure generated-code build quality, performance, security, observability, testability and maintainability alongside translation completeness and manual repair effort.
- Investigate disagreements between tools and oracles, separating a migration defect, an exposed legacy defect and a genuinely ambiguous business rule.
- Deliver the evidence-linked scorecard, sensitivity analysis and acceptance harness so Procurement can defend selection and programme teams can reuse the gates.
Candidate qualifications
- Directed a comparative code-migration or program-transformation benchmark involving mainframe languages and a production target stack.
- Validated semantic equivalence for financial or insurance calculations containing fixed precision, date rules, batch restart and consequential side effects.
- Built behavioural tests where documentation was incomplete, reconciling executable evidence with specialist knowledge and historical production outcomes.
- Controlled AI coding tools and human remediation so productivity claims included prompting, review, repair, testing and operational hardening.
- Assessed translated code for security, performance, observability and long-term maintainability rather than successful compilation alone.
- Produced an independent vendor recommendation that procurement and engineering could reperform from frozen environments, fixtures and retained findings.
Non-negotiables
- The named director must lead benchmark design, vendor controls and final scoring, attend Warsaw reviews and complete the Kraków and Prague execution visits.
- No source code, masked fixture, tool output or vendor prompt history may leave the insurer’s isolated workspaces.
- Will disclose commercial relationships, certification income or referral arrangements involving any migration, cloud or target-platform supplier.
- Has personally benchmarked mainframe transformation with behavioural equivalence evidence; generic software delivery or AI coding-tool adoption is insufficient.
- 49 words maximum. Describe a migrated legacy behaviour that passed ordinary tests but failed semantic equivalence at a precision, boundary or restart condition.
- 49 words maximum. How would you measure human intervention so an AI migration tool cannot hide expert repair behind its automation claim?
- 49 words maximum. Confirm sixteen-week capacity, required travel and any relationship with migration or target-platform vendors in scope.
This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.