Description
As an engineer in this role, you will be primarily focused on analyzing and optimizing the performance of the latest ML models on the latest iPhones and Mac’s. You will work with models created by the most popular ML frameworks (PyTorch, MLX, etc) and will analyze the inference of those models on device to ensure the stack achieves full machine performance on Apple Silicon. The role also includes scripting, coding, and generation of utilities and debug tools to extract, analyze, and report performance and power related metrics for Apple HW. The ideal candidate will have a passion for ML model architectures and ML inference, deep knowledge of GPU and CPU, computer architecture and memory, compilers and HW drivers.
Minimum Qualifications
Experience with ML inference, quantization, performance and accuracy Familiarity and experience with the most popular ML architectures (e.g. LLM’s, Diffusion models, CNN’s) A passion to explore and learn about the latest advances in ML model design and architecture, particularly as related to model implementation on HW and on-device inference Familiarity with Operating Systems, embedded systems, and CPU/GPU HW architectures Highly proficient in Python/C++ and shell scripting Familiarity with Linux or macOS Exceptional clarity in verbal and written communication, including the ability to present and lead discussions in larger groups
Preferred Qualifications
Masters or PhDs in Computer Science or relevant disciplines. Experience with Apple’s CoreML, MPS Graph, Metal Performance Shader’s or MLX frameworks Experience with any on-device ML stack, such as TFLite, ONNX, ExecuTorch, etc. Experience with any ML authoring framework (PyTorch, TensorFlow, JAX, etc.). Experience with Apple’s App development framework such as Xcode, Swift, Objective-C Experience with any compiler stack (MLIR/LLVM/TVM etc.)
Learn more about this Employer on their Career Site
