Confidential mandate
Autonomous AI Red-Team Evidence Architect
Planned Hiring / New
Autonomous AI Red-Team Evidence Architect mandate in Tel Aviv, Israel · AI Security Testing Platforms
An AI security platform commissions a six-week assurance build to make autonomous red-team findings reproducible, separate model failures from harness artefacts and establish an accepted vulnerability evidence chain.
The mandate
The platform’s attack agents generate persuasive transcripts, yet customer-assurance reviewers cannot replay a material share of high-severity findings because model versions, tool states, target fixtures and intermediate decisions drift between runs. The defined problem is to distinguish genuine target weaknesses from stochastic discovery, evaluator error and harness-created exposure without reducing adversarial testing to deterministic scripts.
The deliverable is an Autonomous Red-Team Evidence System comprising a versioned threat library, experiment manifest, target-fixture standard, run-capture schema, replay protocol, severity rubric, causal triage tree, reviewer workflow and customer-ready evidence packet. It must preserve the agent’s exploration while making consequential claims inspectable and bounded.
Delivery starts on 11 January 2027. Milestone one, due 22 January, is the signed threat-to-evidence map and reproducibility baseline; milestone two on 5 February provides the instrumented harness, negative controls and witnessed replay campaign; milestone three on 19 February supplies the accepted operating system, sampled casebook, training materials and adoption backlog.
Acceptance requires product security to reproduce twenty sampled attack paths within agreed behavioural tolerances, explain every material variance and route each finding to target weakness, model behaviour, tool permission or harness defect. The panel must independently rescore the sample from retained evidence with no critical disagreement; a compelling transcript without configuration lineage cannot be accepted.
The client will provide isolated target applications, authorised model endpoints, frozen agent builds, tool sandboxes, prior findings, customer evidence requirements and named adversarial and product engineers. Testing is limited to expressly authorised fixtures; live-customer attacks, vulnerability disclosure, remediation engineering, model retraining and formal certification remain outside the statement of work.
Why this is external work
The researchers who built the attack agents are measured on novel findings and cannot neutrally decide whether their harness created the claimed exposure. Customer teams need stable evidence, but they lack the adversarial depth to preserve open-ended agent behaviour. An external architect can design a falsifiable middle ground and leave acceptance with the accountable security panel.
What you will own
- Classify the threat library by target boundary, attacker objective, permitted tools, success condition, prohibited action and evidence needed for a defensible claim.
- Instrument prompts, model versions, seeds, tool calls, environment state, intermediate plans, target responses and reviewer interventions without hiding exploratory branches.
- Establish negative and counterfactual controls that reveal vulnerabilities introduced by credentials, fixtures, evaluator hints or unsafe harness permissions.
- Design replay tolerances for stochastic attack paths, separating exact transcript reproduction from repeatable exploitability and equivalent harmful outcomes.
- Build a causal triage workflow that assigns model, agent scaffold, tool, target and evaluation failures before severity reaches a customer report.
- Run a witnessed twenty-case campaign and resolve disagreement between automated judges, adversarial researchers, product security and independent reviewers.
- Transfer the evidence schema, runnable fixtures, sampled casebook, severity decisions and maintenance ownership to the assurance team at final acceptance.
Candidate qualifications
- Designed or independently reviewed autonomous-agent red teaming for foundation models, tool-using systems or AI-enabled security products.
- Reproduced a stochastic exploit or harmful agent trajectory using retained state, equivalent paths and well-defined behavioural tolerances.
- Distinguished target vulnerability from evaluation-harness, permission, judge-model or prompt-induced artefact in a contested security finding.
- Built evidence packets that product security, customers and legal reviewers could inspect without disclosing unnecessary exploit or model information.
- Governed authorised attack boundaries, credential handling and vulnerability escalation while researchers pursued open-ended adversarial behaviour.
- Delivered a repeatable evaluation system whose internal owners could operate and recalibrate after the external engagement closed.
Non-negotiables
- The named architect must attend both Tel Aviv laboratory sessions and the regional customer-assurance rehearsal within the six-week window.
- Testing may touch only written-authorised targets and credentials; no exploratory convenience permits activity against live or third-party systems.
- All attack traces, vulnerabilities and model artefacts remain in the client environment and may not enter public benchmarks or personal tools.
- Current relationships with model providers, security vendors, evaluation suppliers or affected customers must be disclosed before case access.
- 49 words maximum. Describe a high-severity autonomous-agent finding that failed reproduction and the evidence that explained why.
- 49 words maximum. How would you preserve open-ended attack exploration while making exploitability independently testable?
- 49 words maximum. Confirm six-week capacity, required Israel travel and every vendor or model-provider relationship relevant to independent assurance.
This mandate is confidential. The client is named only under a mutual NDA, and your own record is never listed, sold or shown to a company under your name until you release it for this specific mandate.