aijobsdesk35,458 roles · 851 companies
← All roles

Research Engineer

KogResearchHybridFull-timeParis, France· posted 3h ago

About Kog

Kog builds the fastest LLM inference engine on standard datacenter GPUs. Our Kog Inference Engine generates 3,000 output tokens per second per request on a single 8× AMD MI300X node and 2,100 on an 8× NVIDIA H200 node (FP16, batch size 1, no speculative decoding).

We co-design the model architecture and the execution engine together. Our Laneformer model uses Delayed Tensor Parallelism (DTP), a novel architecture that restructures the Transformer dependency graph so inter-GPU communication overlaps with computation rather than blocking it.

We pre-trained a 2B-parameter DTP model on 6T tokens on 256 H100 GPUs.

We are a team of 11 people, including 10 engineers and 5 PhDs.

Test it at playground.kog.ai http://playground.kog.ai. Read the technical details on the Kog Labs blog https://blog.kog.ai.

What you will work on

You will imagine, design, and run experiments to understand how architectural decisions propagate through inference behavior, morph existing open-weight models into architecture variants optimized for speed, and turn findings into measurable gains in generation speed and model quality.

- Design new model architecture variants, including routing strategies, attention mechanisms, and MoE structure, with execution constraints as a first-order design input.

- Extend the Laneformer thesis by exploring inference-aware architectural variants such as DTP, Ladder Residual, and PT-Transformer, and finding what compounds at scale.

- Own the post-training pipeline across fine-tuning, evaluation methodology, and adaptation of existing open-weight models toward architecture variants optimized for inference speed.

- Scale the stack to large MoE models such as DeepSeek v4 and Qwen 3, working through routing, expert parallelism, and communication patterns at inference time.

- Write up findings as research papers, submit them to top venues, and present them at conferences.

- Contribute to building AI agents that will perform architecture research and training experiments autonomously, starting from the research foundations we are building now.

What we look for

- You have designed or changed model architecture, where the structure itself was the object of the work. Showing that work, a paper, a repository, or a thesis, is a requirement to move forward.

- You reason about model design and hardware together, tracing how communication structure and layer dependencies shape inference behavior, with fluency in Transformers and MoE deep enough to weigh trade-offs.

- Stronger signals include inference-aware architectural variants such as DTP, Ladder Residual, or PT-Transformer, and post-training methods such as fine-tuning, preference optimization, or quantization, including at research scale.

- A top engineering school or a PhD with concrete architecture work counts, even without industry experience.

What we offer

- Direct access to AMD and NVIDIA datacenter GPUs from day one

- A team where creativity and technical judgment carry weight and where the people closest to the problem shape the key decisions

- Problems that sit on the critical path of model execution speed and that directly influence what the system can become

- A remote-friendly working model, with one mandatory week per month in our Paris office. Travel and accommodation covered by the company.

- Compensation aligned with top AI research profiles, including equity