What frontier labs are hiring for now
TL;DR
Selected live official postings ask for engineering execution, valid measurement, diagnosis, and communication. Some welcome self-taught AI power users, but none of these examples makes coding competence optional.
This is a dated review of selected official postings retrieved on October 2, 2026, not a census of every vacancy or a guarantee that applications remain open. Selection focused on evaluation, research engineering, scientific AI, safety, and the software infrastructure behind them. Employer statements below are short paraphrases. The learning priorities are this course's interpretation of those statements.
OpenAI
Research Engineer, Frontier Evals & Environments asks for technical fundamentals across ML, engineering, systems, or statistics; hands-on work with model and agent workflows; experiments built from ambiguous behavioral problems; and scalable evaluation. It emphasizes measurement reliability and variance, as well as product behavior. Learning implication: demonstrate a hypothesis, controlled pipeline, run, analysis, and decision, rather than only a score.
Researcher, Frontier Risk Mitigations describes safety evaluations, mitigation research, red-teaming pipelines, and collaboration with biology and cybersecurity experts. Its stated fit criteria include AI-safety and research-engineering experience, Python, and technical education. Learning implication: domain knowledge is valuable, but this research role also requires technical depth and evidence of prior work.
Anthropic
Research Engineer, Post-Training Model Evaluations focuses on trustworthy measurements during production training, live evaluation monitoring, regressions, dashboards, and cross-team recommendations. It asks for strong Python, production systems, and evaluations at scale; statistics and post-training experience are additional strengths. Learning implication: learn to question a metric, diagnose a regression, and operate the pipeline under pressure.
Research Engineer, Life Sciences lists demonstrated LLM training/evaluation, Python, large-scale data pipelines, ambiguity handling, and communication as minimum qualifications. Biological background, containerization, and deeper ML experience appear among preferred qualifications. Learning implication: an MD is useful context, but does not substitute for the listed ML engineering experience.
Research Scientist, Life Sciences combines scientific workflows, benchmarks, agent tools, production Python, model training or fine-tuning, and computational tools used by biologists. Learning implication: a domain portfolio should include working software and model-evaluation evidence, alongside scientific credibility.
xAI career site
The live site and linked postings currently display the employer name SpaceXAI. This manual preserves that displayed name and uses the x.ai career site requested by the learner. It makes no additional claim about corporate history.
Software Engineer - Evals describes datasets, grading schemes, agent failure diagnosis, and evaluation infrastructure. It explicitly lists software proficiency and three or more years actively writing code, while also saying it is open to levels from new graduates to senior engineers. Learning implication: the posting contains a tension worth clarifying with the recruiter; do not interpret “all levels” as “no coding required.”
Member of Technical Staff - Evaluation Infrastructure emphasizes reliable distributed systems, inference and orchestration bottlenecks, resource management, and evaluation signal quality. Learning implication: this is a systems-engineering track, requiring substantially more infrastructure practice than a local scorer.
Exceptional Software Engineer is a broad engineering role with fewer explicit technical details. Learning implication: a sparse posting does not waive competence. Ask which project and assessment apply rather than inventing a detailed stack requirement.
An older enterprise-evaluation URL found in search redirected to a jobs index. It is excluded as a current role source.
Google DeepMind
Research Engineer, Advancing Agent Quality lists agent-workflow and ML experience, software/cloud development, and Python or ML frameworks. Responsibilities include realistic environments, calibrated automated raters, trajectory diagnostics, and datasets or rewards for post-training. Learning implication: know how a task, grader, agent trace, and training signal connect. Validate a judge against human evidence.
Evaluation-focused organizations
METR: Member of Technical Staff, Evaluation Execution emphasizes integrating models into scaffolds, inspecting results, improving evaluation software, debugging unfamiliar systems, project management, and professional communication. Learning implication: execution quality and trustworthy interpretation are both core work products.
METR: Task Development Engineer asks for difficult, well-scoped tasks, solvability checks, baselining, strong engineering experience, and experience with hard agent evaluations. Learning implication: task QA is a substantial technical skill, not simply writing hard questions.
Apollo Research: Research Scientist/Engineer (Evaluations) welcomes self-taught candidates while asking for strong Python engineering, messy-data analysis, concise writing, and effective AI-tool use. It values post-training knowledge and Inspect experience. Learning implication: AI-assisted execution is explicitly relevant, but reliable engineering and analysis remain necessary.
Scale: Machine Learning Research Scientist, Evaluations focuses on benchmark design, diagnosing text and multimodal failures, and connecting errors to post-training interventions. Its preferred background includes advanced ML knowledge, related graduate education, and research publications. Learning implication: this track requires deeper ML research preparation than domain grading alone.
What recurs in this selected sample
The recurring pattern is a complete feedback loop: define behavior, build environments and graders, run systems, diagnose failures, and improve either the model or its supporting system. Python, production quality, empirical analysis, ambiguity handling, and communication recur. This is a qualitative interpretation of selected postings, not a measured prevalence estimate across the job market.
| Requirement family | Learn here | Portfolio evidence |
|---|---|---|
| Engineering and debugging | 2–4, 9–10 | Repair with before/after checks and a review |
| Valid measurement | 5–7, 12 | Defensible task contract and tested grader |
| Agents and environments | 11, 21 | Controlled tool workflow and trace analysis |
| Failure diagnosis | 10, 13 | Reproduced mechanism and targeted intervention |
| Post-training literacy | 27 | Explain training signals and reward failures |
| Domain expertise | 15–19 | Source-adjudicated task with valid limits |
| Production operation | 20–21, 29 | Permissions, recovery, complete run status |
| Communication | 14, 23–24 | A short evidence-based decision report |
Do not learn every infrastructure tool before your first eval. Learn the common loop, then choose a role family and build the evidence that family actually asks for.