Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
Ansible @ 4
BGP @ 4
Bash @ 7
CI/CD @ 4
CUDA @ 4
Ceph @ 3
Communication @ 7
Debugging @ 4
GPU @ 4
Go @ 7
HPC @ 6
InfiniBand @ 4
Kubernetes @ 7
Linux @ 7
NCCL @ 4
NVLink @ 4
Networking @ 4
Observability
Python @ 7
Security @ 4
Slurm @ 7
Stress Testing @ 4
Terraform @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
As a Senior/Staff Forward Deployed Engineer focused on AI Infrastructure, you will work at the frontier of large-scale AI systems, taking complex customer infrastructure programs from requirements to working production environments. You will help bring next-generation NVIDIA GPU systems and Groq's purpose-built inference platform into production.
You will operate across GroqCloud, GroqMetal, GPU and LPX infrastructure, networking, storage, orchestration, observability, and workload layers. This hands-on individual-contributor role involves writing code and automation, troubleshooting across infrastructure layers, and translating ambiguous customer requirements into deployed and validated solutions. Field work will also contribute to reusable tooling, deployment patterns, reference architectures, and improvements to the Groq platform.
Responsibilities
- Own technical execution across complex customer engagements, from discovery and architecture through proofs of concept, demonstrations, deployment, cluster bring-up, validation, acceptance, production readiness, and operational handoff.
- Translate incomplete or ambiguous customer requirements into architectures, implementation plans, test criteria, runbooks, and engineering actions.
- Work hands-on across Linux, bare-metal infrastructure, Kubernetes, Slurm, networking, storage, observability, automation, and Groq platform integrations.
- Support large-scale GPU and LPX deployments, including infrastructure bring-up, cluster health and performance validation, workload testing, benchmarking, failure isolation, and production-readiness evidence.
- Analyze AI training and inference workloads, including concurrency, throughput, latency, data movement, caching, scheduling, and infrastructure bottlenecks.
- Lead technical customer discovery, architecture reviews, demonstrations, and proofs of concept.
- Partner with Networking and Security teams on private connectivity, peering, routing, ingress and egress, load balancing, network policy, access controls, security architecture reviews, and enterprise security diligence.
- Troubleshoot production and pre-production issues across organizational and technical boundaries and drive them to resolution.
- Embed with Platform, Cloud, Infrastructure, or Operations teams when customer deployments require additional engineering, automation, or integration work.
- Build reusable tools, automation, reference architectures, test suites, deployment patterns, documentation, and lessons learned.
- Bring structured customer feedback and field evidence to Product and Engineering to help turn recurring gaps into repeatable platform capabilities.
Requirements
- 4+ years of hands-on experience building, deploying, operating, or troubleshooting cloud infrastructure, AI infrastructure, HPC systems, large-scale platforms, or similarly demanding production environments.
- Strong Linux and distributed-systems fundamentals, with practical experience in Kubernetes, Slurm, bare-metal environments, or comparable infrastructure platforms.
- Technical depth in at least one area such as GPU or accelerator systems, networking, storage, orchestration/platform engineering, or infrastructure reliability, with breadth across adjacent layers.
- Working knowledge of AI training and inference workloads and their effects on compute, networking, storage, scheduling, latency, and throughput.
- Strong Python, Go, Bash, or equivalent scripting and programming skills for diagnostics, automation, deployment tooling, testing, or integrations.
- Demonstrated experience personally debugging and delivering systems rather than operating only at the architecture, project-management, or escalation level.
- Ability to break ambiguous problems into concrete technical actions and drive issues to resolution across multiple teams.
- Strong written and verbal communication skills, including requirements gathering, explanation of technical tradeoffs, and documentation.
- Ability to work in a fast-moving environment where customer requirements, platform capabilities, and implementation details evolve in parallel.
Preferred Qualifications
- Experience at a neocloud, hyperscaler, AI infrastructure provider, HPC environment, frontier AI company, or organization operating large-scale accelerator infrastructure.
- Hands-on experience with NVIDIA GPU infrastructure and CUDA, NCCL, NVLink/NVSwitch, DCGM, GPU Operator, InfiniBand, RoCE, Kubernetes, or Slurm.
- Experience bringing up, qualifying, or operating multi-node GPU clusters, including health checks, burn-in or stress testing, collective-communication testing, performance benchmarking, and acceptance criteria.
- Familiarity with VAST, Weka, Lustre, Ceph, or similar high-performance storage systems.
- Experience with Terraform, Ansible, CI/CD, BMC/Redfish, PXE/iPXE, or related infrastructure automation and lifecycle tooling.
- Prior solutions engineering, sales engineering, solutions architecture, or technical pre-sales experience in cloud, networking, security, AI infrastructure, or data center systems.
- Customer-facing networking experience involving private interconnects and peering, BGP, routing, load balancing, Kubernetes/Cilium network policy, and north-south and east-west traffic design.
- Customer-facing security experience involving access-control architecture, network isolation, enterprise security reviews, and SOC 2 or ISO 27001-style diligence.
- Experience defining or executing technical proofs of concept, reference architectures, cluster acceptance tests, performance benchmarks, migration plans, or production-readiness criteria.
Compensation
Total cash salary, inclusive of potential bonus value, depends on level:
- Staff: $270,400–$318,100 per year
- Senior Staff: $341,400–$401,600 per year
Placement within the range depends on geographic location, experience, skills, and internal compensation standards. The ranges apply to candidates located in the United States. International compensation varies based on local market dynamics. Groq also offers a Long-Term Incentive Program and employee benefits.
Additional Information
This position may require access to technology or information subject to U.S. export control laws and regulations. Candidates must meet applicable citizenship or residency criteria or otherwise qualify for an export license. Groq is an Equal Opportunity Employer and provides reasonable accommodations to qualified individuals with disabilities. All offers are contingent on verification of identity and employment authorization. Groq may use artificial intelligence tools or automated systems during recruiting activities, subject to human review and applicable accommodation and consent requirements.