Network Engineer

USD 150,000-250,000 per year
MIDDLE
✅ On-site

Tech Stack

AI Ansible @ 3 BGP @ 3 Communication @ 6 GPU @ 2 HPC NCCL Networking @ 2 Python @ 3 Terraform @ 3

Details

SpaceXAI is building and operating large-scale networks that underpin training and inference infrastructure, including high-performance and supercompute fabrics connecting GPU clusters, as well as the core, edge, and datacenter networks that keep the infrastructure reachable and reliable.

This is a hands-on engineering role focused on owning production network design, deployment, and operations end to end. The role involves designing and building networks, qualifying platforms, shipping changes safely, and maintaining availability and performance at scale. It is not a NOC technician role or a network-software telemetry/ZTP platform software engineering role.

Travel to Memphis and other build sites may be required for capacity build-outs. Participation in a team on-call rotation is required.

Responsibilities

  • Design, deploy, and operate production datacenter and campus/core networks at scale.
  • Own routing and switching configuration standards using BGP and at least one interior gateway protocol such as OSPF or IS-IS, including change design, peer reviews, and execution.
  • Qualify new network platforms, optics, and topologies; contribute to architecture and capacity planning.
  • Build and improve monitoring, alerting, and operational documentation so issues are identified and resolved quickly.
  • Troubleshoot Layer 2 and Layer 3 incidents end to end, from link flaps and optics through routing and traffic engineering, and drive root-cause analysis and lasting fixes.
  • Automate repetitive network tasks with Python, Ansible, or similar tooling where it reduces toil.
  • Partner with compute, facilities, and software teams during cluster build-outs and maintenance windows.
  • Support high-performance and supercompute network environments, including Ethernet AI/HPC fabrics and RoCE/RDMA-capable designs, as part of the broader network estate. Deep specialist RoCE/NCCL ownership is a plus but is not required.

Requirements

  • Several years of experience designing and/or operating production networks in a datacenter, ISP, cloud, or large enterprise environment.
  • Solid hands-on experience with BGP and at least one interior routing protocol.
  • Working knowledge of TCP/IP, VLANs, EVPN/VXLAN or equivalent datacenter overlays, optics, and high-speed Ethernet.
  • Experience troubleshooting live production network incidents and participating in on-call rotations.
  • Strong written and verbal communication skills, including the ability to produce clear change documentation and incident notes.
  • Willingness to work onsite in Palo Alto.

Preferred Skills And Experience

  • Experience with modern datacenter vendors such as Arista, Cisco, Juniper, and Nvidia/Mellanox.
  • Familiarity with high-performance or supercompute networking, including RoCEv2, congestion control, and GPU cluster fabrics.
  • Production experience with network automation using Python, Ansible, Terraform, or similar tools.
  • Experience with EVPN, leaf-spine architectures, and large-scale Ethernet fabrics.
  • Previous experience supporting rapid datacenter or cluster capacity build-outs.

Compensation And Benefits

  • Base salary: $150,000–$250,000 USD per year.
  • Equity.
  • Comprehensive medical, vision, and dental coverage.
  • Access to a 401(k) retirement plan.
  • Short- and long-term disability insurance.
  • Life insurance.
  • Various other discounts and perks.

More jobs at SpaceXAI

Similar jobs