Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
AWS @ 4
Ansible
Automated Testing @ 3
Azure @ 4
BGP @ 7
CI/CD @ 3
Clos Network @ 7
Communication @ 7
Compliance
Design Patterns
Distributed Systems @ 4
GCP @ 4
GPU @ 4
Git
HPC @ 4
InfiniBand @ 3
Linux @ 7
Machine Learning
Networking @ 7
Observability @ 4
Python @ 4
Security
Software Development @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Groq is building high-performance AI infrastructure designed to make inference fast, predictable, and scalable. The infrastructure team designs and operates systems connecting compute, data, and services at significant scale.
This role will architect, design, validate, and evolve network infrastructure supporting Groq's growing AI compute environments. The engineer will own network architecture and design across high-performance data center and AI infrastructure environments, translating capacity, performance, reliability, and business requirements into scalable network designs that can be deployed and operated consistently.
The role involves collaborating with network operations, data center engineering, compute, platform, security, and infrastructure teams to take designs from concept through validation and production deployment. The successful candidate will establish engineering standards, design patterns, tooling, and technical direction to scale Groq's network infrastructure while maintaining performance, reliability, and operational simplicity.
Responsibilities
- Own the end-to-end network design process for AI training and inference environments.
- Translate customer compute, workload, scale, and tenancy requirements into production-ready architectures.
- Collaborate with software teams building inference and training stacks to align network topology with stack architecture.
- Develop high-level designs (HLDs), low-level designs (LLDs), bills of materials (BOMs), and implementation standards using industry and vendor reference architectures.
- Design spine-leaf/Clos fabrics across front-end, back-end, and storage networks.
- Evaluate and implement BGP, EVPN, VXLAN, ECMP, IPv4/IPv6, and QoS.
- Develop strategies for data center interconnect, WAN connectivity, internet edge, cloud connectivity, and external peering.
- Perform network capacity planning and traffic engineering for high-bandwidth AI and distributed compute workloads.
- Establish standards for network hardware platforms, topology, addressing, routing policy, configuration, and lifecycle management.
- Evaluate switches, optics, cabling architectures, network operating systems, and emerging networking technologies.
- Build lab environments, proofs of concept, and design-validation frameworks before production deployment.
- Define failure scenarios and validate network behavior during link, device, control-plane, and site failures.
- Drive network automation and infrastructure-as-code practices using Python, Ansible, Git, APIs, and CI/CD pipelines.
- Develop automated design validation, configuration generation, compliance, and testing capabilities.
- Establish network observability requirements, including telemetry, performance metrics, alerting, and capacity visibility.
- Lead technical reviews and provide guidance on complex network architecture and design decisions.
- Troubleshoot complex performance and reliability issues spanning network, compute, hardware, and distributed systems.
- Partner with operations teams to ensure designs are supportable, observable, and safe to operate at scale.
- Mentor engineers and improve practices for network architecture, automation, documentation, and engineering.
Requirements
- 8+ years of experience in network engineering, architecture, or infrastructure engineering, including significant experience designing large-scale data center networks.
- Deep understanding of TCP/IP, BGP, and modern data center routing architectures.
- Strong experience designing leaf-spine/Clos network fabrics.
- Experience with EVPN/VXLAN and modern network virtualization architectures.
- Strong understanding of ECMP, routing convergence, failure domains, traffic engineering, and network resiliency.
- Experience designing high-bandwidth Ethernet platforms with 100G, 400G, or higher-speed interfaces.
- Strong knowledge of optical networking fundamentals, transceivers, fiber infrastructure, and data center cabling.
- Experience with network platforms and operating systems from Arista, Cisco, Juniper, NVIDIA, or equivalent vendors.
- Experience with network automation and software development using Python and APIs.
- Familiarity with infrastructure-as-code, source control, automated testing, and CI/CD practices.
- Experience with network telemetry and observability technologies, including streaming telemetry, SNMP, flow data, and time-series monitoring.
- Strong understanding of Linux networking and TCP/IP fundamentals.
- Ability to develop architecture documents, design standards, implementation plans, and technical specifications.
- Ability to analyze complex systems and make architectural decisions involving performance, reliability, cost, and operational tradeoffs.
- Strong communication skills and the ability to influence technical decisions across engineering organizations.
- Experience designing networks for AI/ML clusters, HPC environments, distributed computing systems, or large GPU/accelerator deployments.
- Experience with very large-scale Ethernet fabrics and high-throughput east-west traffic patterns.
- Knowledge of RDMA, Priority Flow Control (PFC), ECN, congestion management, and lossless or near-lossless Ethernet design.
- Familiarity with backend technologies such as InfiniBand, RoCE, or Spectrum-X is strongly preferred.
- Experience designing networks supporting distributed training or large-scale AI inference.
- Familiarity with SONiC or other disaggregated/open networking platforms.
- Experience with multi-site data center architecture and data center interconnect.
- Experience with cloud networking environments such as AWS, GCP, or Azure.
- Familiarity with network source-of-truth and automation platforms such as NetBox, Nautobot, or equivalent systems.
- Experience designing and operating infrastructure at hyperscale or in rapidly growing technology environments.
Location and Work Arrangement
The role is based in one of Groq's three hiring hubs: the Dallas, San Francisco, or New York City area. The person hired must be based in one of these areas. Remote work is available while Groq establishes its local office, with the expectation that the role will transition to onsite once the office opens.
Compensation
The total cash salary range for this position, inclusive of potential bonus value, is $270,400–$401,600. Individual placement depends on geographic location, experience, skills, and alignment with internal compensation standards. The range is specific to candidates located in the United States. Compensation for international candidates varies based on local market dynamics. Groq also offers a Long-Term Incentive (LTI) Program and employee benefits.
Additional Information
This position may require access to technology or information subject to U.S. export control laws and regulations, including the Export Administration Regulations (EAR). Candidates must meet applicable citizenship, residency, protected-individual, export-license, and other export-control eligibility requirements.
Groq is an Equal Opportunity Employer committed to an inclusive workplace and reasonable accommodations. All offers are contingent on verification of identity and employment authorization. Groq may use AI tools or automated systems to assist with recruiting activities, subject to human review and applicable legal requirements.