Head of Engineering

Inferact · San Francisco · FullTime

Apply on company site

Overview

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for a Head of Engineering to build and lead the organization developing the systems that power vLLM and Inferact. This role requires an engineering leader with genuine technical credibility at the inference layer—someone who understands GPU and accelerator performance, inference runtimes, ML systems optimization, and hardware-software co-design deeply enough to earn the trust of exceptional staff-level engineers.

You'll partner closely with the founders to scale a senior-heavy, highly specialized engineering team while preserving the technical rigor, speed, and ownership that made vLLM successful. You'll recruit and develop rare ML systems talent, translate ambitious research and infrastructure work into a focused execution plan, strengthen how teams operate, and help Inferact deliver reliable, high-performance inference across models, hardware, and deployment environments.

Skills and Qualifications

Minimum qualifications:

Preferred qualifications:

Bonus points if you have:

Logistics

Job alert

Get new jobs by email

Save this search and get relevant new jobs when they appear.