- Dataset Lifecycle Management: Oversee the collection, organizing, cleaning, and maintenance of large-scale, high-quality datasets for model training.
- Pipeline Development: Build and maintain scalable data processing pipelines and automated intelligent agents to continuously ingest, clean, and enrich training data.
- Quality & Benchmarking: Define, track, and optimize dataset quality metrics (e.g., diversity, absence of bias) to directly improve ML model performance.
- Annotation & Labeling: Design and manage data annotation workflows, collaborating with domain experts to ensure clear, accurate classification protocols.
- Governance & Compliance: Maintain data provenance, ensure compliance with data governance policies (e.g., GDPR, HIPAA if applicable), and enforce data security measures.
- Education: Bachelor’s or Master’s degree in Computer Science, Data Science, Engineering, or a highly quantitative field.
-
Technical Skills:
- Proficiency in programming languages like Python or SQL.
- Experience with Big Data tools and cloud platforms (e.g., AWS, GCP, BigQuery).
- Familiarity with ML frameworks (e.g., PyTorch, Hugging Face).
- Experience: 3+ years managing large-scale datasets, developing data curation heuristics, and working alongside ML researchers or data scientists.
- Analytical Mindset: Strong problem-solving skills to identify data quality anomalies, address model biases, and establish evaluation frameworks.
- A fun, supportive and engaging environment.
- Infrastructures and computational resources to support your work.
- Opportunity to work on cutting edge technologies with the top talents in the field.
- Opportunity to make a significant impact on the transportation revolution by the means of advancing autonomous driving.
- Competitive compensation package.
- Snacks, lunches, dinners, and fun activities.