GPU Engineer

Kog · Paris, France · FullTime

Apply on company site

About Kog

Kog builds the fastest LLM inference engine on standard datacenter GPUs. Our Kog Inference Engine generates 3,000 output tokens per second per request on a single 8× AMD MI300X node and 2,100 on an 8× NVIDIA H200 node (FP16, batch size 1, no speculative decoding).

The hot path is a monokernel implemented with handwritten CUDA (with PTX inline assembly) on NVIDIA, and HIP (with CDNA ISA inline assembly) on AMD.

We optimize at the low level with engine/kernel/model co-design, using reverse engineering to understand and exploit the details of how the GPU hardware works at the micro level.

We are a team of 11 people, including 10 engineers and 5 PhDs.

Test it at playground.kog.ai. Read the technical details on the Kog Labs blog.

What you will work on

You will perform experiments to understand GPU internals, find creative solutions to accelerate critical computational sections used in LLM inference, and write optimized GPU kernels accordingly. Then test, profile, and optimize again.

What we look for

What we offer

Job alert

Get new jobs by email

Save this search and get relevant new jobs when they appear.