Responsibilities
- Drive system validation strategies for AI and HPC hardware platforms, including AI accelerators, GPU clusters, and high-bandwidth memory subsystems in data center environments.
- Perform hands-on bring-up, characterization, and validation of AI server systems and associated components such as PCIe, NVLink, DRAM, and high-speed networking fabrics.
- Develop and maintain test specifications, validation procedures, and debug guides tailored to AI infrastructure NPI programs.
- Investigate and root-cause complex system failures spanning silicon, firmware, software, and hardware layers in collaboration with cross-functional engineering teams.
- Triage and track hardware and firmware defects through resolution while maintaining forward progress on NPI program milestones.
- Identify gaps in test coverage and contribute improvements to test methodologies, tooling, and automation frameworks across the NPI lifecycle.
- Partner with AI platform and capacity engineering teams to define acceptance criteria and deployment readiness standards for new AI hardware systems.
- Support data collection, analysis, and reporting efforts to surface systemic hardware quality trends.
- Communicate validation status and technical findings to internal engineering teams and external hardware vendors.
- Collaborate with firmware and software teams on hardware-software interface requirements for telemetry, diagnostics, and remote management of AI infrastructure.
Minimum Qualifications
- Currently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience. Degree must be completed prior to joining Meta
- 2+ years of experience in hardware systems engineering, silicon validation, firmware validation, or system-level bring-up for AI servers, GPUs, TPUs, or AI accelerator platforms
- Experience in one or more of the following domains: ASIC bring-up and characterization, board-level debug, firmware validation, or large-scale system validation in data center environments
- Experience developing test specifications, validation procedures, and debug methodologies for complex hardware systems
- Experience with root-cause analysis and troubleshooting of system-level failures across hardware, firmware, and software stacks
- Experience with high-speed interconnects or memory subsystems such as PCIe, NVLink, DDR5, or HBM in the context of AI or HPC system validation
Preferred Qualifications
- Experience integrating lab instrumentation and automation frameworks to support NPI validation workflows
- Proficiency in Linux environments and server system management tools used in data center operations
- Experience with debugging tools for SoCs including JTAG, GDB, or Trace32, and familiarity with common bus protocols such as I2C, SPI, USB, and PCIe
- Experience defining hardware-software interface requirements for telemetry, diagnostics, and out-of-band management in AI infrastructure deployments
$118,000/year to $170,000/year + bonus + equity + benefits
Learn more about this Employer on their Career Site
