Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
Ansible @ 3
BGP @ 3
Communication @ 6
GPU @ 2
HPC
NCCL
Networking @ 2
Python @ 3
Terraform @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
SpaceXAI is building and operating large-scale networks that underpin training and inference infrastructure, including high-performance and supercompute fabrics connecting GPU clusters, as well as the core, edge, and datacenter networks that keep the infrastructure reachable and reliable.
This is a hands-on engineering role focused on production network design, deployment, qualification, and operations. The position is not a NOC technician role or a network-software telemetry/ZTP platform software engineering role. Travel to Memphis and other build sites may be required for capacity build-outs, and the role includes participation in a team on-call rotation.
Responsibilities
- Design, deploy, and operate production datacenter and campus/core networks at scale.
- Own routing and switching configuration standards, including BGP and at least one interior gateway protocol such as OSPF or IS-IS.
- Lead change design, peer reviews, and execution.
- Qualify new network platforms, optics, and topologies; contribute to architecture and capacity planning.
- Build and improve monitoring, alerting, and operational documentation.
- Troubleshoot Layer 2 and Layer 3 incidents end to end, from link flaps and optics through routing and traffic engineering, and drive root-cause analysis and lasting fixes.
- Automate repetitive network tasks with Python, Ansible, or similar tooling where it reduces toil.
- Partner with compute, facilities, and software teams during cluster build-outs and maintenance windows.
- Support high-performance and supercompute network environments, including Ethernet AI/HPC fabrics and RoCE/RDMA-capable designs. Deep specialist RoCE/NCCL ownership is a plus but is not required.
Requirements
- Several years of experience designing and/or operating production networks in a datacenter, ISP, cloud, or large enterprise environment.
- Solid hands-on experience with BGP and at least one interior routing protocol.
- Working knowledge of TCP/IP, VLANs, EVPN/VXLAN or equivalent datacenter overlays, optics, and high-speed Ethernet.
- Experience troubleshooting live production network incidents and participating in on-call rotations.
- Strong written and verbal communication skills, including the ability to produce clear change documentation and incident notes.
- Willingness to work onsite in Dublin.
Preferred Skills and Experience
- Experience with modern datacenter vendors such as Arista, Cisco, Juniper, and Nvidia/Mellanox.
- Familiarity with high-performance or supercompute networking, including RoCEv2, congestion control, and GPU cluster fabrics.
- Production experience with network automation using Python, Ansible, Terraform, or similar tools.
- Experience with EVPN, leaf-spine architectures, and large-scale Ethernet fabrics.
- Prior experience supporting rapid datacenter or cluster capacity build-outs.
Compensation and Benefits
- Base salary: $100,000–$150,000 per year.
- Equity.
- Comprehensive medical, vision, and dental coverage.
- Access to a 401(k) retirement plan.
- Short- and long-term disability insurance.
- Life insurance and various other discounts and perks.
More jobs at SpaceXAI
Network Engineer
SpaceXAI · Palo Alto, United States
USD 150,000-250,000 per year
AI Tutor - Video (Weekend)
SpaceXAI · World
USD 40-75 per hour
Team Lead, Human Data Operations - Post-Training
SpaceXAI · United Kingdom, Indonesia, Ireland, India, Japan, South Korea, Philippines, United States, Dubai, United Arab Emirates, Singapore, Singapore
USD 104,000-156,000 per year
Expert Team Lead, Human Data Operations
SpaceXAI · United Kingdom, Indonesia, Ireland, India, Japan, South Korea, Philippines, United States, Dubai, United Arab Emirates, Singapore, Singapore
USD 104,000-170,400 per year
AI Tutor - Design Specialist
SpaceXAI · World
USD 35-75 per hour
Similar jobs
Senior Site Reliability Engineer, DGX Cloud
Nvidia · Santa Clara, United States
USD 168,000-333,500 per year
Solutions Architect
Nebius · Canada
CAD 235,000-300,000 per year
Senior Software Engineer, Core Infrastructure Services - DGX Cloud
Nvidia · United States
USD 168,000-322,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Bare Metal Infrastructure Engineer
Nebius · United States
USD 85,000-140,000 per year
Senior System Software Engineer - GPU Performance
Nvidia · Santa Clara, United States
USD 152,000-287,500 per year
Senior Solution Engineer, Networking
Nvidia · Santa Clara, United States
USD 168,000-322,000 per year