Job Responsibilities:
- Research and develop predictive world models that learn how the physical world evolves, forecasting the future state of a scene from large-scale multimodal driving and robotics data.
- Develop high-quality multi-view future prediction and generation, supporting both action-conditioned rollouts and formulations that forecast the future without explicit action conditioning.
- Work at the boundary between world modeling and policy learning: develop architectures in which a shared backbone both predicts the future and produces trajectories or actions, and apply predictive pre-training to improve Vision-Language-Action (VLA) driving performance.
- Extend prediction beyond 2D pixel into a shared multimodal latent space that spans 3D scene representations such as Gaussian Splatting, together with occupancy and reward signals, so that a single model can support simulation, evaluation, and policy training.
- Advance cross-embodiment generalization: design unified observation and action representations, together with embodiment-conditioning mechanisms, so that a single world model transfers across vehicles, robots, and sensor configurations with only few-shot data.
- Define and build the evaluation methodology for predictive world models, spanning representation quality, prediction accuracy, generation fidelity, physical plausibility, long-horizon rollout consistency, and ultimately closed-loop policy performance, then feed the resulting models back into training as a source of synthetic data and corner-case simulation.
Minimum Skill Requirements:
- MS or PhD level education in Engineering or Computer Science with a focus on Deep Learning, Computer Vision, Generative Models, or a related field, or equivalent experience. Open to recent graduates.
- Strong experience in applied deep learning including model architecture design, large-scale model training, data curation, and empirical analysis.
- 1-3 years + of experience working with DL frameworks such as PyTorch, including hands-on experience with distributed training (FSDP, DeepSpeed, or Megatron-style parallelism).
- Strong Python programming experience with software design skills.
- Solid understanding of data structures, algorithms, code optimization and large-scale data processing.
- Excellent problem-solving skills, including the ability to design controlled experiments and draw sound conclusions from noisy training signals.
Preferred Skill Requirements:
- Hands on experience with generative models for video or 3D, such as diffusion, flow matching, autoregressive video prediction, or neural scene representations including NeRF and Gaussian Splatting.
- Experience with world models or learned simulators for decision making, including model-based reinforcement learning and Vision-Language-Action (VLA) models.
- Experience with multimodal foundation models and video tokenizers or VAEs, including pretraining or adapting large pretrained backbones.
- Experience with large-scale training infrastructure and performance optimization, such as mixed precision, torch.compile, kernel-level optimization, and multi-node scaling.
- A fun, supportive and engaging environment.
- Infrastructures and computational resources to support your work.
- Opportunity to work on cutting edge technologies with the top talents in the field.
- Opportunity to make a significant impact on the transportation revolution by the means of advancing autonomous driving.
- Competitive compensation package.
- Snacks, lunches, dinners, and fun activities.