Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
API @ 3
Ansible @ 3
Bash @ 3
Ceph @ 3
Communication @ 6
HPC @ 3
InfiniBand @ 6
Linux @ 6
Networking @ 3
Python @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Groq is building high-performance AI infrastructure designed to make inference fast, predictable, and scalable. The Network Engineering and Datacenter teams design and operate the systems that provide the compute, storage, and networking capabilities behind Groq's global AI infrastructure.
The Storage Engineer, Deployment & Support will deploy, validate, and operate storage infrastructure supporting global AI and HPC workloads. The role owns storage system bring-up, configuration, validation, performance acceptance, troubleshooting, expansion, upgrades, and ongoing operational health across large-scale environments. It takes approved storage designs from implementation planning through production deployment and operational handoff, partnering closely with Platform Engineering to ensure storage capabilities are reliable, production-ready, and consumable.
Responsibilities
- Deploy and validate storage infrastructure across new and existing Groq datacenters, including high-performance storage systems supporting AI/HPC workloads.
- Configure and bring up storage systems against approved designs and implementation standards, including cluster initialization, storage nodes, interfaces, protocols, and supporting infrastructure.
- Drive storage deployment execution end to end, directing and coordinating on-site datacenter technicians and vendors through physical rack/stack, cabling, and installation while owning configuration, validation, troubleshooting, performance acceptance, and production handoff.
- Develop and maintain implementation procedures, validation plans, deployment checklists, system inventories, as-built documentation, and operational runbooks.
- Perform storage acceptance and performance validation, including throughput, IOPS, latency, capacity, data distribution, cluster health, and recovery behavior.
- Monitor and maintain storage health, capacity, availability, and performance, identifying and resolving issues before they impact production workloads.
- Troubleshoot storage and hardware failures across storage nodes, drives, network paths, protocols, and supporting infrastructure, and perform root cause analysis for deployment and production issues.
- Execute and coordinate storage expansions, upgrades, hardware refreshes, and other lifecycle activities while minimizing production impact.
- Manage hardware readiness for deployments and expansions, including rack/stack coordination, inventory and asset record updates, RMA coordination, spares, and vendor shipments.
- Coordinate deployment audits and execute established QA/QC processes to ensure storage infrastructure meets Groq standards.
- Partner closely with storage design, network, compute, and platform teams to identify and resolve blockers throughout deployment and operations.
- Provide operational support during and after deployments, including maintenance activities, incidents, remediation, break-fix events, and vendor escalation when required.
- Develop tools and scripts that improve storage deployment, configuration, validation, testing, health checks, and repeatable operational workflows.
- Capture lessons learned from deployments and incidents and contribute to global implementation standards and deployment playbooks.
Requirements
- 4+ years of experience in storage engineering, systems engineering, network engineering, datacenter operations, or related infrastructure roles, with hands-on experience deploying, operating, or troubleshooting storage systems.
- Strong hands-on experience operating and troubleshooting large-scale distributed or high-performance storage environments.
- Experience with modern storage platforms and technologies such as VAST, WEKA, DDN, Lustre, Ceph, or comparable distributed storage solutions.
- Strong Linux systems administration and troubleshooting skills in production environments.
- Strong understanding of storage protocols and data paths, including NFS, NFS over RDMA, NVMe-oF, and high-throughput Ethernet or InfiniBand connectivity.
- Strong understanding of storage performance concepts, including throughput, IOPS, latency, capacity, and utilization.
- Understanding of RDMA, RoCE, InfiniBand, and other high-performance networking concepts used in AI/HPC storage environments.
- Ability to troubleshoot systematically across storage, Linux, network, and hardware layers and drive issues through root cause and resolution.
- Experience developing deployment or operational tooling using Python, Bash, Ansible, APIs, or similar technologies.
- Strong ownership, communication, and time-management skills, with the ability to manage multiple deployments and operational priorities under demanding timelines.
- Comfortable operating in fast-moving environments where deployment plans and requirements can evolve quickly.
- Ability to travel to global datacenter locations for deployments, maintenance activities, and other site-specific needs.
Compensation
The total cash salary range for this position, inclusive of potential bonus value, is $270,400–$401,600. Individual placement is determined by geographic location, experience, skills, and alignment with internal compensation standards. This range is specific to candidates located in the United States. Compensation for international candidates varies based on local market dynamics. Groq also offers a Long-Term Incentive Program and employee benefits.
Additional Information
This position may require access to technology or information subject to U.S. export control laws and regulations, including the Export Administration Regulations. Candidates may need to meet applicable citizenship, residency, or export license eligibility criteria. Groq is an Equal Opportunity Employer and provides reasonable accommodations to qualified applicants. All offers are contingent upon verification of identity and employment authorization. The company may use artificial intelligence tools or automated systems to assist with recruiting activities, subject to human review.