Kubernetes Platform Engineer

at Groq
USD 223,600-365,800 per year
SENIOR
✅ Remote ✅ On-site

Tech Stack

AI Distributed Systems @ 7 Kubernetes @ 7 Networking @ 6 Observability

Details

Groq is building a global cloud platform for AI inference workloads, powered by its LPU technology. This role will own and evolve the Kubernetes platform and related hyperscaler technologies that underpin Groq's cloud infrastructure.

Responsibilities

  • Own the architecture, development, reliability, and evolution of Groq's Kubernetes-based platform.
  • Design and build Kubernetes-based systems that operate reliably and efficiently at scale.
  • Build software, automation, and platform capabilities that simplify how workloads are deployed, operated, and scaled.
  • Improve the resilience, observability, performance, and operational efficiency of the Kubernetes platform.
  • Identify and eliminate reliability and scalability bottlenecks across clusters and supporting infrastructure.
  • Own platform capabilities through their full lifecycle, from design and implementation through production operations and continuous improvement.
  • Partner with software and infrastructure engineers to understand workload requirements and translate them into durable platform capabilities.
  • Establish and improve Kubernetes engineering patterns, tooling, and operational practices.
  • Debug complex production issues spanning Kubernetes, distributed systems, networking, compute, and cloud infrastructure.
  • Help shape the technical direction of the platform as its scale and requirements evolve.

Requirements

  • Strong software engineering fundamentals with experience building and operating production infrastructure.
  • Deep hands-on experience with Kubernetes and proficiency designing Kubernetes-based architectures.
  • Experience owning Kubernetes environments in production, including deployment, scaling, upgrades, reliability, and troubleshooting.
  • Experience building software and automation for cloud or infrastructure platforms.
  • Strong understanding of distributed systems and the reliability and scalability challenges of operating infrastructure at scale.
  • Experience with cloud infrastructure and modern infrastructure-as-code and automation practices.
  • Ability to debug complex problems across application, Kubernetes, networking, compute, and infrastructure layers.
  • Comfortable taking end-to-end ownership of production systems and making pragmatic architectural trade-offs.
  • Strong collaboration skills and the ability to translate engineering teams' needs into scalable platform solutions.
  • Motivation to take ownership, pursue technical excellence, improve reliability, and build durable infrastructure.

Location and Work Policy

The role is based in one of Groq's hiring hubs in the Dallas, San Francisco, or New York City area. Employees may work remotely while the local Groq office is established, with the expectation that the role will transition to onsite work once the office opens.

Compensation and Benefits

The total cash salary range is $223,600–$365,800, inclusive of potential bonus value. Individual placement depends on geographic location, experience, skills, and internal compensation standards. The range applies to candidates located in the United States; international compensation varies based on local market dynamics. Groq also offers a Long-Term Incentive Program and employee benefits.

This position may require access to technology or information subject to U.S. export control laws and regulations. U.S. candidates must meet applicable citizenship, residency, protected-individual, or export-license eligibility requirements. International candidates must meet relevant export-control and local-law requirements.

More jobs at Groq

Similar jobs