Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
AWS @ 4
Ansible
Automated Testing @ 3
Azure @ 4
BGP @ 7
CI/CD @ 3
Clos Network @ 7
Communication @ 4
Compliance
Design Patterns
Distributed Systems @ 4
GCP @ 4
GPU @ 4
Git
HPC @ 4
InfiniBand @ 3
Linux @ 7
Machine Learning
Networking @ 7
Observability @ 4
Python @ 4
Security
Software Development @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Groq is building high-performance AI infrastructure designed to make inference fast, predictable, and scalable. The infrastructure teams design and operate systems connecting compute, data, and services at significant scale.
This role will architect, design, validate, and evolve network infrastructure supporting Groq's growing AI compute environments. The engineer will own network architecture and design across high-performance data center and AI infrastructure environments, translating capacity, performance, reliability, and business requirements into scalable network designs that can be deployed and operated consistently.
The role will work closely with network operations, data center engineering, compute, platform, security, and infrastructure teams to take designs from concept through validation and production deployment. It will also help establish engineering standards, design patterns, tooling, and technical direction for scaling Groq's network infrastructure.
Responsibilities
- Own the end-to-end network design process for AI training and inference, translating compute, workload, scale, and tenancy requirements into production-ready architectures.
- Align network topology with software stacks used for inference and training.
- Develop high-level designs (HLDs), low-level designs (LLDs), bills of materials (BOMs), and implementation standards using industry and vendor reference architectures.
- Design spine-leaf/Clos fabrics across front-end, back-end, and storage networks, focusing on performance, resiliency, scalability, and operational simplicity.
- Evaluate and implement BGP, EVPN, VXLAN, ECMP, IPv4/IPv6, and QoS.
- Develop strategies for data center interconnect, WAN connectivity, internet edge, cloud connectivity, and external peering.
- Perform network capacity planning and traffic engineering for high-bandwidth AI and distributed compute workloads.
- Partner with compute and infrastructure engineering teams to understand workload communication patterns and translate them into network requirements.
- Establish standards for network hardware platforms, topology, addressing, routing policy, configuration, and lifecycle management.
- Evaluate switches, optics, cabling architectures, network operating systems, and emerging networking technologies.
- Build lab environments, proofs of concept, and design-validation frameworks before introducing architectures into production.
- Define failure scenarios and validate network behavior during link, device, control-plane, and site failures.
- Drive network automation and infrastructure-as-code using Python, Ansible, Git, APIs, and CI/CD pipelines.
- Develop automated design validation, configuration generation, compliance, and testing capabilities.
- Establish network observability requirements, including telemetry, performance metrics, alerting, and capacity visibility.
- Lead technical reviews and provide guidance on complex network architecture and design decisions.
- Troubleshoot complex performance and reliability issues spanning network, compute, hardware, and distributed systems.
- Partner with operations teams to ensure designs are supportable, observable, and safe to operate at scale.
- Mentor engineers and raise the technical bar for network architecture, automation, documentation, and engineering practices.
Requirements
- 8+ years of experience in network engineering, architecture, or infrastructure engineering, with significant experience designing large-scale data center networks.
- Deep understanding of TCP/IP, BGP, and modern data center routing architectures.
- Strong experience designing leaf-spine/Clos network fabrics.
- Experience with EVPN/VXLAN and modern network virtualization architectures.
- Strong understanding of ECMP, routing convergence, failure domains, traffic engineering, and network resiliency.
- Experience designing networks with high-bandwidth Ethernet platforms, including 100G, 400G, or higher-speed interfaces.
- Strong knowledge of optical networking fundamentals, transceivers, fiber infrastructure, and data center cabling.
- Experience with network platforms and operating systems from Arista, Cisco, Juniper, NVIDIA, or equivalent vendors.
- Experience with network automation and software development using Python and APIs.
- Familiarity with infrastructure-as-code, source control, automated testing, and CI/CD practices.
- Experience with network telemetry and observability technologies such as streaming telemetry, SNMP, flow data, and time-series monitoring.
- Strong understanding of Linux networking and TCP/IP fundamentals.
- Ability to develop architecture documents, design standards, implementation plans, and technical specifications.
- Ability to analyze complex systems and make architectural decisions involving performance, reliability, cost, and operational tradeoffs.
- Strong communication skills and the ability to influence technical decisions across engineering organizations.
- Experience designing networks for AI/ML clusters, HPC environments, distributed computing systems, or large GPU/accelerator deployments.
- Experience with very large-scale Ethernet fabrics and high-throughput east-west traffic patterns.
- Knowledge of RDMA, Priority Flow Control (PFC), ECN, congestion management, and lossless or near-lossless Ethernet design.
- Familiarity with InfiniBand, RoCE, or Spectrum-X is strongly preferred.
- Experience designing networks supporting distributed training or large-scale AI inference.
- Familiarity with SONiC or other disaggregated/open networking platforms.
- Experience with multi-site data center architecture and data center interconnect.
- Experience with cloud networking environments such as AWS, GCP, or Azure.
- Familiarity with network source-of-truth and automation platforms such as NetBox, Nautobot, or equivalent systems.
- Experience designing and operating infrastructure at hyperscale or in rapidly growing technology environments.
Compensation
The total cash salary range is $270,400-$401,600, inclusive of potential bonus value. Individual placement depends on geographic location, experience, skills, and alignment with internal compensation standards. This range applies to candidates located in the United States. International compensation will vary based on local market dynamics. Groq also offers a Long-Term Incentive (LTI) Program and employee benefits.
Work Location
The role is based in one of Groq's three hiring hubs: the Dallas, San Francisco, or New York City area. The person hired must be based in one of these areas. The role may initially be performed remotely while Groq establishes its local office, with the expectation that it will transition to onsite work once the office opens.
Additional Information
The position may require access to technology or information subject to U.S. export control laws and regulations, including the Export Administration Regulations (EAR). Candidates must meet applicable citizenship, residency, or export-license eligibility criteria. Groq is an Equal Opportunity Employer committed to an inclusive environment and provides reasonable accommodations to qualified individuals with disabilities. All offers are contingent upon verification of identity and employment authorization.