Research Engineer

Kog · Paris, France · FullTime

Apply on company site

About Kog

Kog builds the fastest LLM inference engine on standard datacenter GPUs. Our Kog Inference Engine generates 3,000 output tokens per second per request on a single 8× AMD MI300X node and 2,100 on an 8× NVIDIA H200 node (FP16, batch size 1, no speculative decoding).

We co-design the model architecture and the execution engine together. Our Laneformer model uses Delayed Tensor Parallelism (DTP), a novel architecture that restructures the Transformer dependency graph so inter-GPU communication overlaps with computation rather than blocking it.

We pre-trained a 2B-parameter DTP model on 6T tokens on 256 H100 GPUs.

We are a team of 11 people, including 10 engineers and 5 PhDs.

Test it at playground.kog.ai. Read the technical details on the Kog Labs blog.

What you will work on

You will imagine, design, and run experiments to understand how architectural decisions propagate through inference behavior, morph existing open-weight models into architecture variants optimized for speed, and turn findings into measurable gains in generation speed and model quality.

What we look for

What we offer

Job alert

Get new jobs by email

Save this search and get relevant new jobs when they appear.