Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 3
API @ 3
Ansible @ 3
Bash @ 3
Ceph @ 3
Communication @ 6
HPC @ 3
InfiniBand @ 6
Linux @ 6
Networking @ 3
Python @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Groq is building high-performance AI infrastructure designed to make inference fast, predictable, and scalable. This role is responsible for deploying, validating, and operating the storage infrastructure behind Groq's global AI infrastructure across large-scale AI/HPC environments.
The Storage Engineer, Deployment & Support owns the storage infrastructure lifecycle from implementation planning through production deployment and operational handoff. The role includes storage system bring-up, configuration, validation, performance acceptance, troubleshooting, expansion, upgrades, and ongoing operational health. This position partners closely with Platform Engineering, storage design, network, compute, and other infrastructure teams to ensure storage capabilities are reliable, production-ready, and consumable.
Responsibilities
- Deploy and validate storage infrastructure across new and existing Groq datacenters, including high-performance storage systems supporting AI/HPC workloads.
- Configure and bring up storage systems against approved designs and implementation standards, including cluster initialization, storage nodes, interfaces, protocols, and supporting infrastructure.
- Direct and coordinate on-site datacenter technicians and vendors through physical rack/stack, cabling, and installation while owning configuration, validation, troubleshooting, performance acceptance, and production handoff.
- Develop and maintain implementation procedures, validation plans, deployment checklists, system inventories, as-built documentation, and operational runbooks.
- Perform storage acceptance and performance validation, including throughput, IOPS, latency, capacity, data distribution, cluster health, and recovery behavior.
- Monitor and maintain storage health, capacity, availability, and performance, identifying and resolving issues before they impact production workloads.
- Troubleshoot storage and hardware failures across storage nodes, drives, network paths, protocols, and supporting infrastructure, and perform root cause analysis.
- Execute and coordinate storage expansions, upgrades, hardware refreshes, and other lifecycle activities while minimizing production impact.
- Manage hardware readiness for deployments and expansions, including rack/stack coordination, inventory and asset record updates, RMA coordination, spares, and vendor shipments.
- Coordinate deployment audits and execute established QA/QC processes to ensure storage infrastructure meets Groq standards.
- Provide operational support during and after deployments, including maintenance activities, incidents, remediation, break-fix events, and vendor escalation.
- Develop tools and scripts that improve storage deployment, configuration, validation, testing, health checks, and repeatable operational workflows.
- Capture lessons learned from deployments and incidents and contribute to global implementation standards and deployment playbooks.
- Travel to global datacenter locations for deployments, maintenance activities, and other site-specific needs.
Requirements
- 4+ years of experience in storage engineering, systems engineering, network engineering, datacenter operations, or related infrastructure roles, with hands-on experience deploying, operating, or troubleshooting storage systems.
- Strong hands-on experience operating and troubleshooting large-scale distributed or high-performance storage environments.
- Experience with modern storage platforms and technologies such as VAST, WEKA, DDN, Lustre, Ceph, or comparable distributed storage solutions.
- Strong Linux systems administration and troubleshooting skills in production environments.
- Strong understanding of storage protocols and data paths, including NFS, NFS over RDMA, NVMe-oF, and high-throughput Ethernet or InfiniBand connectivity.
- Strong understanding of storage performance concepts, including throughput, IOPS, latency, capacity, and utilization.
- Understanding of RDMA, RoCE, InfiniBand, and other high-performance networking concepts used in AI/HPC storage environments.
- Ability to troubleshoot systematically across storage, Linux, network, and hardware layers and drive issues through root cause and resolution.
- Experience developing deployment or operational tooling using Python, Bash, Ansible, APIs, or similar technologies.
- Strong ownership, communication, and time-management skills, with the ability to manage multiple deployments and operational priorities under demanding timelines.
- Ability to operate in fast-moving environments where deployment plans and requirements can evolve quickly.
Compensation
The total cash salary range for this position, inclusive of potential bonus value, is $270,400–$401,600. Individual placement is determined by geographic location, experience, skills, and alignment with internal compensation standards. This range is specific to candidates located in the United States. Compensation for international candidates varies based on local market dynamics. Groq also offers a Long-Term Incentive Program and employee benefits.
This position may require access to technology or information subject to U.S. export control laws and regulations, including the Export Administration Regulations. Candidates must meet applicable citizenship, residency, or export-license eligibility criteria. Groq is an equal opportunity employer and provides reasonable accommodations to qualified individuals.