Senior/Staff Forward Deployed Engineer, AI Infrastructure

at Groq
USD 270,400-401,600 per year
SENIOR
✅ On-site

Tech Stack

AI @ 6 Ansible @ 4 BGP @ 4 Bash @ 7 CI/CD @ 4 CUDA @ 4 Ceph @ 3 Communication @ 7 Debugging @ 4 GPU @ 4 Go @ 7 HPC @ 6 InfiniBand @ 4 Kubernetes @ 7 Linux @ 7 NCCL @ 4 NVLink @ 4 Networking @ 4 Observability Python @ 7 Security @ 4 Slurm @ 7 Stress Testing @ 4 Terraform @ 4

Details

As a Senior/Staff Forward Deployed Engineer focused on AI Infrastructure, you will work at the frontier of large-scale AI systems, taking complex customer infrastructure programs from requirements to working production environments. You will help bring next-generation NVIDIA GPU systems and Groq’s purpose-built inference platform into production.

You will operate across GroqCloud, GroqMetal, GPU and LPX infrastructure, networking, storage, orchestration, observability, and workload layers. This is a hands-on individual-contributor role involving systems work, coding and automation, troubleshooting across the stack, and translating ambiguous customer requirements into deployed and validated solutions. Field work will also contribute to reusable tooling, deployment patterns, reference architectures, and improvements to the Groq platform.

Responsibilities

  • Own technical execution across complex customer engagements, from discovery and architecture through proofs of concept, demonstrations, deployment, cluster bring-up, validation, acceptance, production readiness, and operational handoff.
  • Translate incomplete or ambiguous customer requirements into architectures, implementation plans, test criteria, runbooks, and engineering actions.
  • Work hands-on across Linux, bare-metal infrastructure, Kubernetes, Slurm, networking, storage, observability, automation, and Groq platform integrations.
  • Support large-scale GPU and LPX deployments, including infrastructure bring-up, cluster health and performance validation, workload testing, benchmarking, failure isolation, and production-readiness evidence.
  • Reason about AI training and inference workloads, concurrency, throughput, latency, data movement, caching, scheduling, and infrastructure bottlenecks.
  • Lead technical customer discovery, architecture reviews, demonstrations, and proofs of concept, explaining design choices, tradeoffs, performance results, and risks.
  • Partner with Networking and Security teams on private connectivity and peering, routing, ingress and egress, load balancing, network policy, access controls, security architecture reviews, and enterprise security diligence.
  • Troubleshoot production and pre-production issues across organizational and technical boundaries and drive them to resolution.
  • Embed with Platform, Cloud, Infrastructure, or Operations teams when customer deployments require engineering execution, automation, or integration work.
  • Build reusable tools, automation, reference architectures, test suites, deployment patterns, documentation, and lessons learned.
  • Bring structured customer feedback and field evidence to Product and Engineering to help turn recurring gaps into repeatable platform capabilities.

Requirements

  • 4+ years of hands-on experience building, deploying, operating, or troubleshooting cloud infrastructure, AI infrastructure, HPC systems, large-scale platforms, or similarly demanding production environments.
  • Strong Linux and distributed-systems fundamentals, with practical experience in Kubernetes, Slurm, bare-metal environments, or comparable infrastructure platforms.
  • Technical depth in at least one area such as GPU or accelerator systems, networking, storage, orchestration and platform engineering, or infrastructure reliability.
  • Working knowledge of AI training and inference workloads and their effects on compute, networking, storage, scheduling, latency, and throughput.
  • Strong Python, Go, Bash, or equivalent scripting and programming skills for diagnostics, automation, deployment tooling, testing, or integrations.
  • Demonstrated experience personally debugging and delivering systems.
  • Ability to break ambiguous problems into concrete technical actions and drive issues to resolution across multiple teams.
  • Strong written and verbal communication skills, including requirements gathering, technical explanation, and documentation.
  • Comfort working in a fast-moving environment with evolving customer requirements and platform capabilities.

Preferred Qualifications

  • Experience at a neocloud, hyperscaler, AI infrastructure provider, HPC environment, frontier AI company, or organization operating large-scale accelerator infrastructure.
  • Experience with NVIDIA GPU infrastructure and CUDA, NCCL, NVLink/NVSwitch, DCGM, GPU Operator, InfiniBand, RoCE, Kubernetes, or Slurm.
  • Experience bringing up, qualifying, or operating multi-node GPU clusters, including health checks, burn-in or stress testing, collective-communication testing, performance benchmarking, and acceptance criteria.
  • Familiarity with VAST, Weka, Lustre, Ceph, or similar high-performance storage technologies.
  • Experience with Terraform, Ansible, CI/CD, BMC/Redfish, PXE/iPXE, or related infrastructure lifecycle tooling.
  • Prior solutions engineering, sales engineering, solutions architecture, or technical pre-sales experience in cloud, networking, security, AI infrastructure, or data center systems.
  • Customer-facing networking experience with private interconnects and peering, BGP and routing, load balancing, Kubernetes/Cilium network policy, and north-south and east-west traffic design.
  • Customer-facing security experience with access-control architecture, network isolation, enterprise security reviews, and SOC 2 or ISO 27001-style diligence.
  • Experience defining or executing technical proofs of concept, reference architectures, cluster acceptance tests, performance benchmarks, migration plans, or production-readiness criteria.

Compensation

Groq uses a Total Cash philosophy that incorporates potential bonus value directly into base pay. The total cash salary ranges for this position are dependent on level:

  • Staff: $270,400–$318,100 per year
  • Senior Staff: $341,400–$401,600 per year

Individual placement depends on geographic location, experience, skills, and alignment with internal compensation standards. The listed ranges apply to candidates located in the United States. International compensation will vary based on local market dynamics. Groq also offers a Long-Term Incentive Program and employee benefits.

Additional Information

This position may require access to technology or information subject to U.S. export control laws and regulations. Candidates must qualify as U.S. Persons for export control purposes or otherwise be eligible for an applicable export license. Groq is an Equal Opportunity Employer and provides reasonable accommodations to qualified individuals with disabilities.

More jobs at Groq

Similar jobs