Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
AWS @ 4
Ansible
Automated Testing @ 3
Azure @ 4
BGP @ 7
CI/CD @ 3
Clos Network @ 7
Communication @ 7
Compliance
Design Patterns
Distributed Systems @ 4
GCP @ 4
GPU @ 4
Git
HPC @ 4
InfiniBand @ 4
Linux @ 7
Machine Learning
Networking @ 7
Observability @ 4
Python @ 4
Security
Software Development @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Groq is building high-performance AI infrastructure to make inference fast, predictable, and scalable. The infrastructure team designs and operates systems connecting compute, data, and services at significant scale.
This role will architect, design, validate, and evolve network infrastructure supporting Groq's AI compute environments. The engineer will own network architecture and design across high-performance data center and AI infrastructure environments, translating capacity, performance, reliability, and business requirements into scalable, repeatable network designs.
The role collaborates with network operations, data center engineering, compute, platform, security, and infrastructure teams to take designs from concept through validation and production deployment. It also establishes engineering standards, design patterns, tooling, and technical direction for scaling Groq's network infrastructure.
Responsibilities
- Own the end-to-end network design process for AI training and inference environments.
- Translate customer compute, workload, scale, and tenancy requirements into production-ready architectures.
- Work with software teams building inference and training stacks to align network topology with stack architecture.
- Develop high-level designs (HLDs), low-level designs (LLDs), bills of materials (BOMs), and implementation standards using industry and vendor reference architectures.
- Design spine-leaf/Clos fabrics across front-end, back-end, and storage networks.
- Evaluate and implement BGP, EVPN, VXLAN, ECMP, IPv4/IPv6, and QoS.
- Develop strategies for data center interconnect, WAN connectivity, internet edge, cloud connectivity, and external peering.
- Perform capacity planning and traffic engineering for high-bandwidth AI and distributed compute workloads.
- Establish standards for network hardware platforms, topology, addressing, routing policy, configuration, and lifecycle management.
- Evaluate switches, optics, cabling architectures, network operating systems, and emerging networking technologies.
- Build lab environments, proofs of concept, and design-validation frameworks.
- Define and validate failure scenarios involving links, devices, control planes, and sites.
- Drive network automation and infrastructure-as-code using Python, Ansible, Git, APIs, and CI/CD pipelines.
- Develop automated design validation, configuration generation, compliance, and testing capabilities.
- Establish network observability requirements, including telemetry, performance metrics, alerting, and capacity visibility.
- Lead technical reviews and guide complex network architecture and design decisions.
- Troubleshoot performance and reliability issues spanning networking, compute, hardware, and distributed systems.
- Partner with operations teams to ensure designs are supportable, observable, and safe to operate at scale.
- Mentor engineers and improve practices for network architecture, automation, documentation, and engineering.
Requirements
- 8+ years of experience in network engineering, architecture, or infrastructure engineering, including significant large-scale data center network design.
- Deep understanding of TCP/IP, BGP, and modern data center routing architectures.
- Strong experience designing leaf-spine/Clos network fabrics and EVPN/VXLAN network virtualization architectures.
- Strong understanding of ECMP, routing convergence, failure domains, traffic engineering, and network resiliency.
- Experience designing high-bandwidth Ethernet platforms with 100G, 400G, or higher-speed interfaces.
- Strong knowledge of optical networking fundamentals, transceivers, fiber infrastructure, and data center cabling.
- Experience with network platforms and operating systems from Arista, Cisco, Juniper, NVIDIA, or equivalent vendors.
- Experience with network automation and software development using Python and APIs.
- Familiarity with infrastructure-as-code, source control, automated testing, and CI/CD practices.
- Experience with network telemetry and observability technologies, including streaming telemetry, SNMP, flow data, and time-series monitoring.
- Strong understanding of Linux networking and TCP/IP fundamentals.
- Ability to develop architecture documents, design standards, implementation plans, and technical specifications.
- Ability to make architectural decisions involving performance, reliability, cost, and operational tradeoffs.
- Strong communication skills and ability to influence technical decisions across engineering organizations.
- Experience designing networks for AI/ML clusters, HPC environments, distributed computing systems, or large GPU/accelerator deployments.
- Experience with very large-scale Ethernet fabrics and high-throughput east-west traffic.
- Knowledge of RDMA, Priority Flow Control (PFC), ECN, congestion management, and lossless or near-lossless Ethernet design.
- Backend networking experience with InfiniBand, RoCE, or Spectrum-X is strongly preferred.
- Experience designing networks for distributed training or large-scale AI inference.
- Familiarity with SONiC or other disaggregated/open networking platforms.
- Experience with multi-site data center architecture and data center interconnect.
- Experience with cloud networking environments such as AWS, GCP, or Azure.
- Familiarity with network source-of-truth and automation platforms such as NetBox, Nautobot, or equivalent systems.
- Experience designing and operating infrastructure at hyperscale or in rapidly growing technology environments.
Compensation
The total cash salary range is $270,400-$401,600, inclusive of potential bonus value. Individual placement depends on geographic location, experience, skills, and alignment with internal compensation standards. This range applies to candidates located in the United States. International compensation varies by local market. Groq also offers a Long-Term Incentive (LTI) Program and employee benefits.
Location and Work Policy
The role is based in one of Groq's three hiring hubs: the Dallas, San Francisco, or New York City area. Employees may work remotely while Groq establishes a local office, with the expectation that the role will transition to onsite once the office opens.
Additional Information
The position may require access to technology or information subject to U.S. export control laws and regulations. Candidates must meet applicable citizenship, residency, protected-person, or export-license eligibility requirements. Groq is an equal opportunity employer and provides reasonable accommodations to qualified applicants.