- Optimize transformer-based LLMs for low-latency and high-throughput inference.
- Optimize kernels and model graphs using tools like CUDA, Triton, and custom fused operators.
- Implement and benchmark (Quantization, Knowledge distillation, structured and unstructured pruning, KV-cache optimization, etc.).
- Deploy optimized models across GPUs, CPUs, and edge acceleators.
- Contribute to internal tooling and documentation for model optimization flows.
- Master in CS/CE/EE, or equivalent, with 3 + years of industry experience.
- Good knowledge of PyTorch.
- Knowledge of transformer architecture and ways to accelerate the training and inference of transformer models.
- Previous experience in the autonomous driving industry.
- Knowledge of Torchscript and Nvidia TensorRT.
- Strong programming skills in Python and C++
- Familiarity with GPU CPU, NPU, DSP architecture.
- Deep understanding of memory bandwidth, compute bottlenecks, and hardware-aware model optimization
- Being efficiently in solving complex problems collaboratively on larger teams
- A fun, supportive and engaging environment.
- Infrastructures and computational resources to support your work.
- Opportunity to work on cutting edge technologies with the top talents in the field.
- Opportunity to make significant impact on the transportation revolution by the means of advancing autonomous driving.
- Competitive compensation package.
- Snacks, lunches, dinners, and fun activities.