About the Role
Together AI is looking for a Senior Network Engineer to design, deploy, and operate the global network infrastructure supporting our production services and high-performance AI compute environments.
This is a hands-on engineering role for someone with deep networking expertise who can also troubleshoot across Linux, Kubernetes, automation, and application boundaries. You will work on large-scale, multi-vendor data center networks and help ensure they remain highly available, reliable, scalable, and performant.
The ideal candidate has strong networking fundamentals, experience operating complex networks at scale, and a structured, evidence-based approach to troubleshooting. You should be comfortable owning problems from initial investigation through root cause and resolution, including situations where the issue may extend beyond the network itself.
Requirements
- 8+ years of professional experience designing, building, and supporting large-scale production data center, cloud, service-provider, or high-performance computing networks (excluding enterprise networks).
- Deep understanding of TCP/IP and strong experience with technologies such as BGP, OSPF, VXLAN, EVPN, ECMP, and QoS.
- Experience designing and supporting multi-tenant network environments using technologies such as VRFs, VLANs, overlays, and policy-based segmentation.
- Hands-on experience deploying and troubleshooting network platforms from vendors such as Arista, Cisco, Juniper, and NVIDIA.
- Strong troubleshooting skills using tools such as Wireshark, tcpdump, MTR, curl, nmap, and standard Linux networking utilities.
- Ability to diagnose connectivity, latency, packet-loss, routing, and performance issues across the network, host, and application layers.
- Experience developing or maintaining network automation using Python, Ansible, or similar tools.
- Experience working through a Git-based software development lifecycle, including branching, code review, validation, linting, testing, CI/CD, deployment, and rollback.
- Working knowledge of Kubernetes networking, including pods, services, CNIs, and basic connectivity troubleshooting.
- Foundational knowledge of RDMA networking and technologies such as RoCE or InfiniBand.
- Experience with cloud networking in AWS, GCP, or Azure.
- Strong Linux administration and troubleshooting skills.
Responsibilities
- Design, deploy, operate, and maintain global, multi-vendor, multi-protocol networks supporting high-performance AI compute infrastructure.
- Troubleshoot complex network and application-connectivity issues, identify root causes, and drive problems through resolution.
- Analyze telemetry, packet captures, logs, and performance data to identify network degradation, congestion, packet loss, and capacity constraints.
- Participate in architecture and design reviews to ensure solutions meet requirements for performance, availability, scalability, security, and operational supportability.
- Develop and maintain automation, validation, and operational tooling that improves network reliability and reduces manual effort.
- Evaluate network hardware, software, optics, and emerging technologies for use in production environments.
- Establish standards and operational best practices for network design, deployment, monitoring, change management, and incident response.
- Lead projects addressing complex technical challenges and contribute directly to the network engineering roadmap.
- Partner with infrastructure, systems, security, and application teams to troubleshoot issues that cross traditional ownership boundaries.
Preferred
- Hands-on experience deploying or operating RoCE and/or InfiniBand fabrics.
- Experience supporting GPU clusters, HPC environments, distributed storage, or other high-bandwidth and latency-sensitive workloads.
- Understanding of AI training and inference traffic patterns and the demands they place on network infrastructure.
- Experience operating networks spanning thousands of devices, multiple data centers, and multiple geographic regions.
- Familiarity with AI-assisted engineering tools and the ability to validate, test, and safely deploy AI-generated automation or code.
About Together AI
Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month.
Compensation
We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $190,000 - $280,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.
Equal Opportunity
Please see our privacy policy at https://www.together.ai/privacy