Software Engineer, ML Networking

USD 315,000-425,000 per year
MIDDLE
✅ Hybrid

SCRAPED

Used Tools & Technologies

Not specified

Required Skills & Competences ?

Algorithms @ 3 Data Structures @ 6 Distributed Systems @ 3 Communication @ 6 Networking @ 3 Performance Optimization @ 3 Rust @ 3 Debugging @ 3

Details

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

Role summary

A systems-level engineer specializing in network infrastructure and network optimization, with expertise in building and maintaining software that interacts with networks. You will be responsible for writing and maintaining software that interfaces between our accelerators and our high-speed networks. This role requires deep technical knowledge of network protocols, kernel-space and/or user-space networks, interfacing with hardware, and the ability to debug and optimize distributed software at the network level.

Responsibilities

  • Write and maintain software that interfaces between ML accelerators and high-speed networks.
  • Diagnose and resolve networking issues in distributed systems (especially OSI layers 2-4).
  • Build higher-level abstractions such as collectives and RPC.
  • Benchmark and optimize networking and OS-level performance for large-scale workloads.
  • Implement and optimize collective algorithms and congestion control for synchronous workloads.
  • Work on projects such as accelerator-initiated tensor movement over the network, benchmarking new networking environments, implementing new collective algorithms to improve latency, and debugging kernel-level network latency spikes.

Requirements

Networking Systems Engineering

  • Expert-level proficiency with network protocols and networking concepts.
  • Deep kernel networking knowledge: TCP/IP stack internals, XDP, eBPF, io_uring, and epoll.
  • User-space networking experience: DPDK, RDMA, and kernel bypass techniques.
  • Understanding of how to build higher-level abstractions like collectives and RPC.
  • Skilled at diagnosing and resolving networking issues in distributed systems, especially at OSI layers 2-4.

Low-Level Systems and OS Programming

  • Strong programming skills in systems programming languages with attention to memory management, lock-free data structures, and NUMA-aware programming.
  • Experience with software, driver, and OS performance optimization tools and techniques.
  • Comfort with or desire to learn Rust.

Strong candidates may have

  • Understanding of ML accelerators and accelerator drivers.
  • Demonstrated ability to design new network protocols.
  • Experience with PCIe and drivers for PCIe devices.
  • Expertise in algorithms used in networking, including compression and graph algorithms.
  • Experience programming on SmartNICs.
  • 5+ years of experience in systems programming or network programming.
  • Backgrounds such as HPC, telecommunications, host networking software, OS/kernel engineering, or embedded systems.
  • Strong debugging mindset and patience for complex, multi-layered issues.

Representative projects

  • Build a system for accelerator-initiated tensor movement over the network.
  • Benchmark software for a new networking environment.
  • Implement a new collective algorithm to improve latency.
  • Optimize congestion control algorithms for large-scale synchronous workloads.
  • Debug kernel-level network latency spikes.

Logistics & compensation

  • Annual salary: $315,000 - $425,000 USD.
  • Education: At least a Bachelor's degree in a related field or equivalent experience.
  • Location-based hybrid policy: staff are expected to be in one of our offices at least 25% of the time (some roles may require more time in office).
  • Visa sponsorship: Anthropic does sponsor visas and retains immigration counsel to assist where possible.

Benefits & culture

Anthropic is a public benefit corporation headquartered in San Francisco. They offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and collaborative office space. The team emphasizes large-scale, high-impact AI research, strong communication, and diverse perspectives. Candidates are encouraged to apply even if they don't meet every listed qualification.