Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
API @ 3
Ansible @ 3
Bash @ 3
Ceph @ 3
Communication @ 6
HPC @ 3
InfiniBand @ 6
Linux @ 6
Networking @ 3
Python @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Groq is building high-performance AI infrastructure designed to make inference fast, predictable, and scalable. The Network Engineering and Datacenter teams design and operate the systems that provide the compute, storage, and networking capabilities behind Groq's AI infrastructure.
This role is responsible for deploying, validating, and operating storage infrastructure supporting global AI and HPC environments. The Storage Engineer will take approved storage designs from implementation planning through production deployment and operational handoff, owning storage system bring-up, configuration, validation, performance acceptance, troubleshooting, expansion, upgrades, and ongoing operational health.
Responsibilities
- Deploy and validate storage infrastructure across new and existing Groq datacenters, including high-performance storage systems supporting AI/HPC workloads.
- Configure and bring up storage systems against approved designs and implementation standards, including cluster initialization, storage nodes, interfaces, protocols, and supporting infrastructure.
- Direct and coordinate on-site datacenter technicians and vendors through physical rack/stack, cabling, and installation while owning configuration, validation, troubleshooting, performance acceptance, and production handoff.
- Develop and maintain implementation procedures, validation plans, deployment checklists, system inventories, as-built documentation, and operational runbooks.
- Perform storage acceptance and performance validation, including throughput, IOPS, latency, capacity, data distribution, cluster health, and recovery behavior.
- Monitor and maintain storage health, capacity, availability, and performance, identifying and resolving issues before they impact production workloads.
- Troubleshoot storage and hardware failures across storage nodes, drives, network paths, protocols, and supporting infrastructure, and perform root cause analysis.
- Execute and coordinate storage expansions, upgrades, hardware refreshes, and other lifecycle activities while minimizing production impact.
- Manage hardware readiness for deployments and expansions, including rack/stack coordination, inventory and asset record updates, RMA coordination, spares, and vendor shipments.
- Coordinate deployment audits and execute established QA/QC processes to ensure storage infrastructure meets Groq standards.
- Partner with storage design, network, compute, and platform teams to identify and resolve deployment and operational blockers.
- Provide operational support during and after deployments, including maintenance activities, incidents, remediation, break-fix events, and vendor escalation.
- Develop tools and scripts to improve storage deployment, configuration, validation, testing, health checks, and repeatable operational workflows.
- Capture lessons learned from deployments and incidents and contribute to global implementation standards and deployment playbooks.
Requirements
- 4+ years of experience in storage engineering, systems engineering, network engineering, datacenter operations, or related infrastructure roles, with hands-on experience deploying, operating, or troubleshooting storage systems.
- Strong hands-on experience operating and troubleshooting large-scale distributed or high-performance storage environments.
- Experience with modern storage platforms and technologies such as VAST, WEKA, DDN, Lustre, Ceph, or comparable distributed storage solutions.
- Strong Linux systems administration and troubleshooting skills in production environments.
- Strong understanding of storage protocols and data paths, including NFS, NFS over RDMA, NVMe-oF, and high-throughput Ethernet or InfiniBand connectivity.
- Strong understanding of storage performance concepts, including throughput, IOPS, latency, capacity, and utilization.
- Understanding of RDMA, RoCE, InfiniBand, and other high-performance networking concepts used in AI/HPC storage environments.
- Ability to troubleshoot systematically across storage, Linux, network, and hardware layers and drive issues through root cause and resolution.
- Experience developing deployment or operational tooling using Python, Bash, Ansible, APIs, or similar technologies.
- Strong ownership, communication, and time-management skills, with the ability to manage multiple deployments and operational priorities under demanding timelines.
- Comfortable operating in fast-moving environments where deployment plans and requirements can evolve quickly.
- Ability to travel to global datacenter locations for deployments, maintenance activities, and other site-specific needs.
Compensation
The total cash salary range for this position, inclusive of potential bonus value, is $270,400–$401,600 for candidates located in the United States. Individual placement depends on geographic location, experience, skills, and alignment with internal compensation standards. Compensation for international candidates varies based on local market dynamics. Groq also offers a Long-Term Incentive Program and employee benefits.
Additional Information
This position may require access to technology or information subject to U.S. export control laws and regulations, including the Export Administration Regulations. Candidates may need to meet applicable citizenship, residency, or export-license eligibility criteria. Groq is an Equal Opportunity Employer and provides reasonable accommodations to qualified individuals with disabilities.