SonicJobs Logo
Left arrow iconBack to search

Member of Technical Staff, Head of Quality

Plato
Posted 8 days ago, valid for 19 days
Location

San Francisco, CA, US

Salary

$180,000 - $280,000 per year

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • Plato is seeking a Head of Quality with 3+ years of experience in software engineering, ML engineering, or research systems to build and lead their QA systems from scratch.
  • The role involves owning the final sign-off on environments and datasets, ensuring correctness and signal density while automating verification infrastructure.
  • Candidates should have a strong understanding of adversarial systems and be able to identify potential loopholes in reward functions and task designs.
  • The position offers a competitive salary and the opportunity to make a significant impact in the field of reinforcement learning by improving the quality of training signals.
  • Plato, based in San Francisco, is backed by leading investors and researchers and aims to transform the way AI agents are trained.

About Plato

Plato is an applied research lab building the environments used to train specialized AI agents. We turn proprietary real-world data into high-fidelity simulations that produce the dense reinforcement learning signal required to train frontier models.

Compute and baseline architectures are rapidly commoditizing; environment design and RL data are the true bottlenecks. Today, models don’t fail from a lack of compute—they fail because bad task design, leaky reward functions, and brittle verifiers train them to cheat instead of learn. Plato exists to make RL signal provable, grounded, and robust at scale.

We’re based in San Francisco and backed by leading investors and researchers across top frontier labs.

Why Apply

  • Own the Ground Truth: You will own the final checkpoint between raw environments and customer model runs. If an environment teaches a model the wrong behavior, you pull the plug.

  • Massive Leverage: We don’t solve quality by throwing armies of manual labelers at a spreadsheet. You will architect the automated judge harnesses, red-teaming agents, and telemetry that enforce quality programmatically.

  • Founding Impact: As Head of Quality (Member of Technical Staff), you will build Plato’s QA systems from scratch, set the technical bar, and scale the team.

  • Frontier Signal: Work directly on the failure modes frontier labs face when scaling post-training, test-time compute, and agentic workflows.

The Role

In RL, quality is not a polish step—it is the training signal itself.

When task designs are ambiguous, sandboxes lack proper isolation, or reward functions contain loopholes, agents do what optimizers always do: exploit the grader. A broken environment doesn’t just burn cluster hours; it actively poisons downstream model weights.

As Head of Quality, you will own the standard for what constitutes real learning signal across every task, verifier, and environment Plato ships. You’ll be the adversarial mind finding the loopholes before the model does, the engineer automating the test suites, and the leader directing the team running verification.

What You’ll Do

  • Gate Final Delivery: Own the final sign-off before environments and datasets ship to frontier labs, auditing trajectories, tasks, and reward dynamics for correctness, feasibility, and signal density.

  • Harden Verifiers & Tasks: Build automated red-teaming suites to stress-test task feasibility and verifier integrity, aggressively eliminating reward hacking, grader tampering, and impossible task traps.

  • Automate Verification Infra: Architect judge models, sandbox replay harnesses, and rollout forensics to catch synthetic drift, out-of-distribution behaviors, and leaky states upstream.

  • Build & Lead the Verification Team: Hire and direct a high-agency team of QA engineers, domain specialists, and technical reviewers, blending automated agentic checks with deep human-in-the-loop review.

  • Close the Loop with Research: Translate downstream model failure modes into concrete generator constraints so defects are prevented at the generation stage rather than caught in review.

Required Qualifications

  • Technical Depth in ML/RL or Systems: 3+ years of experience in software engineering, ML engineering, or research systems, with proficiency in modern languages (Python, etc.).

  • Adversarial Systems Instinct: You understand how optimizers exploit edges. You naturally think about how an agent could game a reward function, break a sandbox, or fake completion.

  • Trajectory & Code Forensics: Proven ability to dive into raw rollouts, agent reasoning traces, and verification code to spot subtle ungrounded assumptions or hallucinated logic.

  • High-Agency Leadership: Experience building or leading a technical evaluation, QA, or data verification function from zero to one.

  • Zero-Compromise Bar: Comfort holding the line on delivery under intense customer pressure. You understand that shipping contaminated data is far worse than shipping late.

Preferred Qualifications

  • Experience with RL training dynamics, automated LLM evals, agent sandboxing, or synthetic trajectory generation.

  • Background designing adversarial test suites, code execution verifiers, or formal verification systems.

  • Experience handling client-facing technical evaluations and failure postmortems with frontier AI research teams.Introduction

Plato is an applied research lab building the foundational infrastructure to train specialized AI agents.

We turn real-world data streams into high-fidelity simulated environments that generate the training signal needed to make capable models. Our work supports frontier labs, hyperscalers, and enterprises building AI systems for complex, high-stakes work.

Today, only a handful of players can train models for capable work. Compute and algorithms are rapidly commoditizing, but reinforcement learning data remains the bottleneck. Plato is changing that by automatically scaling training environments from proprietary real-world data.




Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.