Senior/Staff Forward Deployed Engineer, AI Infrastructure

at Groq
USD 270,400-401,600 per year
SENIOR
✅ On-site

Tech Stack

AI @ 6 Ansible @ 4 BGP @ 4 Bash @ 7 CI/CD @ 4 CUDA @ 4 Ceph @ 3 Communication @ 7 GPU @ 4 Go @ 7 HPC @ 6 InfiniBand @ 4 Kubernetes @ 7 Linux @ 7 NCCL @ 4 NVLink @ 4 Networking @ 4 Observability Python @ 7 Security @ 4 Slurm @ 7 Stress Testing @ 4 Terraform @ 4

Details

As a Senior/Staff Forward Deployed Engineer focused on AI Infrastructure, you will work at the frontier of large-scale AI systems, taking complex customer infrastructure programs from requirements to working production environments. You will help bring accelerator and infrastructure technologies into production, spanning next-generation NVIDIA GPU systems and Groq’s purpose-built inference platform.

You will operate across GroqCloud, GroqMetal, GPU and LPX infrastructure, and the networking, storage, orchestration, observability, and workload layers around them. This is a hands-on individual-contributor role involving systems work, code and automation, troubleshooting across infrastructure layers, and converting ambiguous customer requirements into deployed and validated solutions. Field work will also inform reusable tooling, deployment patterns, reference architectures, and improvements to the Groq platform.

Responsibilities

  • Own technical execution across complex customer engagements, from discovery and architecture through proofs of concept, demonstrations, deployment, cluster bring-up, validation, acceptance, production readiness, and operational handoff.
  • Translate incomplete or ambiguous customer requirements into practical architectures, implementation plans, test criteria, runbooks, and concrete engineering actions.
  • Work hands-on across Linux, bare-metal infrastructure, Kubernetes, Slurm, networking, storage, observability, automation, and Groq platform integrations.
  • Support large-scale GPU and LPX deployments, including infrastructure bring-up, cluster health and performance validation, workload testing, benchmarking, failure isolation, and production-readiness evidence.
  • Analyze AI training and inference workloads, including concurrency, throughput, latency, data movement, caching, scheduling, and infrastructure bottlenecks.
  • Lead technical customer discovery, architecture reviews, demonstrations, and proofs of concept, explaining design choices, tradeoffs, performance results, and risks to engineering and business stakeholders.
  • Partner with Networking and Security teams on private connectivity and peering, routing, ingress and egress, load balancing, network policy, access controls, security architecture reviews, and enterprise security diligence.
  • Troubleshoot production and pre-production issues across organizational and technical boundaries, driving them to resolution while coordinating internal experts.
  • Embed with Platform, Cloud, Infrastructure, or Operations teams when customer deployments require concentrated engineering execution, automation, or integration work.
  • Build reusable tools, automation, reference architectures, test suites, deployment patterns, documentation, and lessons learned.
  • Provide structured customer feedback and field evidence to Product and Engineering to help turn one-off solutions into repeatable platform capabilities.

Requirements

  • 4+ years of hands-on experience building, deploying, operating, or troubleshooting cloud infrastructure, AI infrastructure, HPC systems, large-scale platforms, or similarly demanding production environments.
  • Strong Linux and distributed-systems fundamentals, with practical experience in Kubernetes, Slurm, bare-metal environments, or comparable infrastructure platforms.
  • Technical depth in at least one area such as GPU or accelerator systems, networking, storage, orchestration/platform engineering, or infrastructure reliability, with breadth across adjacent layers.
  • Working knowledge of AI training and inference workloads and their effects on compute, networking, storage, scheduling, latency, and throughput.
  • Strong Python, Go, Bash, or equivalent scripting and programming skills for diagnostics, automation, deployment tooling, testing, or integrations.
  • Demonstrated ability to personally debug and deliver systems rather than operate only at the architecture, project-management, or escalation level.
  • Ability to break ambiguous problems into concrete technical actions and drive issues to resolution across multiple teams.
  • Strong written and verbal communication skills, including requirements gathering, explanation of technical tradeoffs, and technical documentation.
  • Comfort working in a fast-moving environment where customer requirements, platform capabilities, and implementation details evolve in parallel.

Preferred Qualifications

  • Experience at a neocloud, hyperscaler, AI infrastructure provider, HPC environment, frontier AI company, or other organization operating large-scale accelerator infrastructure.
  • Hands-on experience with NVIDIA GPU infrastructure and CUDA, NCCL, NVLink/NVSwitch, DCGM, GPU Operator, InfiniBand, RoCE, Kubernetes, or Slurm.
  • Experience bringing up, qualifying, or operating multi-node GPU clusters, including health checks, burn-in or stress testing, collective-communication testing, performance benchmarking, and acceptance criteria.
  • Familiarity with VAST, Weka, Lustre, Ceph, or similar high-performance storage systems and distributed AI workload data-access patterns.
  • Experience with Terraform, Ansible, CI/CD, BMC/Redfish, PXE/iPXE, or related infrastructure automation and lifecycle tooling.
  • Prior solutions engineering, sales engineering, solutions architecture, or technical pre-sales experience in cloud, networking, security, AI infrastructure, or data center systems.
  • Customer-facing networking experience involving private interconnects and peering, BGP and routing, load balancing, Kubernetes/Cilium network policy, and north-south and east-west traffic design.
  • Customer-facing security experience involving access-control architecture, network isolation, enterprise security reviews, and SOC 2 or ISO 27001-style diligence.
  • Experience defining or executing technical proofs of concept, reference architectures, cluster acceptance tests, performance benchmarks, migration plans, or production-readiness criteria.

Compensation

  • Staff total cash salary: $270,400–$318,100 per year.
  • Senior Staff total cash salary: $341,400–$401,600 per year.
  • Total cash compensation includes potential bonus value directly in base pay. Individual placement depends on geographic location, experience, skills, and internal compensation standards.
  • The listed ranges apply to candidates located in the United States. International compensation varies based on local market dynamics.
  • Groq also offers a Long-Term Incentive Program and employee benefits.

Additional Information

This position may require access to technology or information subject to U.S. export control laws and regulations, including the Export Administration Regulations. Candidates must meet applicable citizenship or residency criteria or qualify for an applicable export license. Groq is an Equal Opportunity Employer and provides reasonable accommodations. All offers are contingent upon verification of identity and employment authorization.

More jobs at Groq

Similar jobs