Responsibilities
- Design and implement scalable systems for distributed ML training and inference, including data ingestion pipelines, feature processing, and model serving infrastructure
- Develop and optimize ML platform components such as training orchestration, gradient communication, and checkpoint management across large-scale distributed environments
- Profile and diagnose performance bottlenecks across the ML stack, including data loading, preprocessing, forward and backward passes, and serving latency
- Write automated tests covering expected behaviors, failure modes, and error paths for ML systems components, and build monitoring and alerting for production anomalies
- Collaborate with ML researchers and product engineers to translate model requirements into reliable, high-throughput system designs
- Own technical design for features and components within ML infrastructure, evaluating trade-offs between throughput, latency, cost, and engineering maintainability
- Participate in staged rollouts of ML system changes using feature flagging and A/B testing frameworks, monitoring key metrics and responding to regressions
- Contribute to code quality through code reviews, clear technical documentation, and consolidation of duplicative implementations across the ML systems codebase
- Support on-call rotations for owned ML infrastructure, investigating production incidents and contributing detailed retrospectives to prevent recurrence
Minimum Qualifications
- Currently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Degree must be completed prior to joining Meta
- 2+ years of experience in software engineering with a focus on machine learning systems, distributed systems, or high-performance data infrastructure
- Experience writing production-quality code in Python and at least one compiled language such as C++ or Java, including performance-sensitive systems code
- Experience designing and implementing components of ML training or inference pipelines, including data preprocessing, model execution, or serving systems
- Experience with distributed computing concepts such as data parallelism, model parallelism, or parameter server architectures as applied to ML workloads
- Experience writing automated tests, building logging and alerting, and participating in production incident response for large-scale systems
Preferred Qualifications
- Experience optimizing ML workloads for throughput and latency, including profiling GPU or CPU utilization, memory bandwidth, and communication overhead in distributed training or inference settings
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
- Experience with ML framework internals such as PyTorch or JAX, including custom operator development, execution graph optimization, or compiler integration
- Experience with stream or batch data processing systems such as Apache Spark, Flink, or Kafka in the context of ML feature pipelines
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
- 1+ years production experience in GenAI post-training, RLHF, RLVR, comms/collectives, parallelism, accuracy evaluation, and performance profiling
- Familiarity with model quantization, mixed-precision training, or other techniques for reducing compute and memory costs in production ML systems
$121,992/year to $181,000/year + bonus + equity + benefits
Learn more about this Employer on their Career Site
