Founding AI Research Lead - Agentic AI Lab
San Francisco Bay Area | Full time
Backed by 8VC, we are building a world-class team to tackle one of industry's most critical problems: trusted AI for enterprise operations.
About the Role
Fabrion is designing the future of enterprise AI infrastructure, grounded in agents, knowledge graphs, and multi-tenant governance. We are working on research inside the Agentic AI Lab to train and evaluate specialized models for mission-critical enterprise work.
The direction is specific and ambitious. We share the full thesis under NDA during the interview process. What we can say here: the program has committed design partners with production data access, dedicated compute, a benchmark-first plan with clear go and no-go gates, and a platform team that has already built the governance and serving layer your models will run behind.
This is full-cycle research: problem formulation, data, training, evaluation, and deployment, with your name on the results.
Core Responsibilities
Own the research agenda: model and training design, evaluation protocol, and the publication plan
Take models from public benchmark results to live customer shadow deployments, with gates you define and defend
Set the benchmark discipline: strong baselines first, published comparables cited, results that survive scrutiny
Lead and grow a small team (ML engineer, data engineer, contractors) and pair closely with the founders and platform team
Write technical plans internally and papers externally when results warrant it
Desired Experience
Hands-on experience training sequence models, owning the tokenizer, the training loop, and the evaluation, not only fine-tuning through APIs
Strong background in at least two of: reinforcement learning (especially offline and imitation settings), sequence decision modeling, structured or constrained generation, learning from event and log data
A track record of shipping research into a product or landing a rigorous benchmark result
PhD in machine learning or a closely related field, or an equivalent research record
Preferred Tech Stack
PyTorch, the Hugging Face ecosystem, experiment tracking and reproducible training pipelines, modern cloud data warehouses, evaluation harness engineering
Soft Skills & Mindset
Comfortable as the most senior researcher in the room: setting direction under ambiguity and writing decisions down
Rigor over hype: you distrust your own results until the baselines agree
A teacher's instinct: part of this role is turning strong engineers into researchers
Why This Role Matters
We believe specialized models built on governed enterprise data can run real, multi-billion-dollar workflows. Your work will not be buried in research reports. It will be benchmarked in public, deployed to real customers, and activated by hundreds of thousands of decisions.
Learn more about this Employer on their Career Site
