Overview
About METR
Full job description
About METR We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation and misalignment. We believe it is robustly good for policymakers and civil society to have a clear understanding of risks from AI systems, and we are extremely excited to build a team of ambitious, excellent people to tackle one of the most important challenges of our time.
METR has started embedding researchers inside frontier labs to investigate incidents, stress-test labs’ internal agent monitoring systems, and assess loss-of-control risks from internal deployment. As agent capabilities increase, we expect this to be one of the most important sources of independent information the world has about catastrophic risks from advanced AI.
Recent incidents have involved complex multi-day cyber attacks on frontier lab internal infrastructure and external third parties. As we further develop our incident investigation and embedded stress-testing capacity, we will need talented cyberforensics researchers who can conduct embedded exercises. We expect these assessors to have deep access, and for their work to be a large part of METR's impact in the next year. We want to build on the momentum from previous exercises to further develop our risk assessments.
Incident investigation: You'll be embedded in a frontier AI lab for up to several weeks at a time, likely alongside 1-4 other METR staff. Between exercises, you'll practice, develop the general methodology, talk to other researchers, build tooling to make future exercises go better, help us hire and scale, write up results, and plan/coordinate future exercises.
Red-teaming: You will attack agent monitoring and security systems, potentially embedded in labs or red-teaming METR internal infrastructure.
Reporting: You'd produce findings rigorous enough for lab boards, governments, and the public and contribute to METR's public incident tracking and risk reports.
Building AI-assisted forensic tooling: Incidents at our scale (tens of thousands of actions) often can't be read solely by hand. You'd build LLM-powered pipelines to triage transcripts, cluster behaviors, flag deception, and accelerate future investigations.
Digital forensics and incident response: You have investigated severe security incidents end to end. You have experience with evidence acquisition and preservation, log and timeline reconstruction across cloud, network, endpoint, and identity systems, attacker tradecraft analysis, and post-incident reporting.
Cloud and infrastructure fluency: You can follow an intrusion through AWS (CloudTrail, IAM, VPC flow logs), Kubernetes and containers, CI/CD, and package registries.
Understanding LLMs: You know how frontier models are trained and deployed (RL post-training, agent scaffolds, sandboxing, monitoring) well enough to reason about root causes, and you build and analyze with LLMs.
Attention to detail and communication: You can run rigorous investigations and write findings that hold up to scrutiny.
Experience investigating incidents involving AI agents, or research on agent misbehavior, deception, or sandbox escapes.
Exploit and vulnerability analysis.
Experience with training-data analysis, model internals/interpretability, or running experiments on model checkpoints.
Formal investigation experience: NTSB/CSB-style safety investigations, law enforcement or intelligence forensics, regulatory or expert-witness work.
Familiarity with the tooling in our environment: DataDog, Kubernetes, CrowdStrike Falcon, Okta, Tailscale, Pulumi, PostgreSQL.
Tips for this job
Practical Job and Scholarship guidance. These tips do not replace official rules or create new eligibility requirements.
- Tailor the CV and application to the responsibilities and required skills stated on the official employer page.
- Use concrete evidence of relevant work, projects and measurable results rather than generic claims.
- Confirm location, work authorization, remote restrictions and sponsorship terms before applying.
- Apply through the original employer or official recruitment destination shown on this page.
Verification notes
laptop-ats-crawler v2
Job and Scholarship is the discovery and verification layer. Confirm eligibility, dates, salary/funding and application instructions on the original source before submitting anything.
Metr (lever) ↗Browse current Job and Scholarship listings from Metr (lever) →