Senior AI Platform Engineer, Infrastructure Services

USD 132,000-182,000 per year
SENIOR
✅ Remote

Tech Stack

AI @ 1 API @ 4 AWS @ 4 Apigee @ 4 ArgoCD @ 7 CI/CD @ 4 Claude Code @ 4 Communication @ 6 DevOps @ 6 GPU @ 4 GitHub @ 4 GitHub Actions Jenkins @ 4 Kubernetes @ 7 LLM @ 1 LLMOps @ 3 Mentoring @ 4 Observability @ 1 RAG @ 4 Security Terraform @ 4 Vector Databases @ 4 vLLM

Details

SentinelOne is seeking a Senior AI Platform Engineer, Infrastructure Services, to own its AI Gateway infrastructure, built on Kong AI Gateway. The platform authenticates, routes, rate-limits, and monitors AI coding assistant traffic across the organization. This high-autonomy role will set technical direction for AI infrastructure, lead incident response and reliability efforts, and design solutions spanning CI/CD, GitOps, Kubernetes, and artifact systems.

Responsibilities

  • Architect, harden, and scale the Kong AI Gateway deployment using Konnect Hybrid on KCP/EKS.
  • Manage authentication with Okta/OIDC, consumer tiers and budgets, rate limiting, semantic caching, and observability.
  • Lead root-cause analysis and remediation for gateway issues involving timeouts, latency, capacity, and failover.
  • Build monitoring and alerting to identify issues before they affect users.
  • Design solutions across Jenkins, JPAAS, ArgoCD, Kubernetes, Artifactory/Xray, GitHub Enterprise, and GitHub Actions runner fleets.
  • Evaluate and roll out AI developer tooling, including Qodo for AI-assisted pull request review and LinearB for engineering metrics.
  • Make build-versus-buy recommendations for AI infrastructure and developer tooling.
  • Define architecture and engineering standards, review designs, and mentor other engineers.
  • Partner with security, DevEx, and product engineering teams to translate platform needs into capabilities.
  • Deploy and operate self-hosted and open-weight model serving infrastructure using technologies such as vLLM, NVIDIA Triton/NIM, TGI, and Ollama.
  • Plan GPU capacity, configure autoscaling, and tune cost and performance for model-serving workloads.
  • Help establish LLMOps practices for model versioning, evaluation, and safe rollout.
  • Support retrieval-augmented generation infrastructure, including vector stores and embedding pipelines.
  • Build observability for token usage, latency, and spending across API-based and self-hosted models.

Requirements

  • At least 5 years of experience in platform, infrastructure, or DevOps engineering, with experience owning production systems end to end.
  • Hands-on experience with API gateways such as Kong, Envoy, Apigee, or similar. Experience with AI/LLM gateway patterns, including rate limiting, semantic caching, and prompt/response observability, is a strong plus.
  • Strong Kubernetes and GitOps experience, including ArgoCD or a comparable platform, across development, government, and production environments.
  • Solid CI/CD experience, including Jenkins pipeline design and administration, build infrastructure, and runner or agent fleet management.
  • Experience with artifact and package management systems such as Artifactory and Xray, or similar tools.
  • Experience administering source control platforms such as GitHub Enterprise.
  • Working knowledge of Terraform and AWS/EKS.
  • Experience deploying and operating self-hosted LLM inference stacks and GPU-backed infrastructure, including Kubernetes GPU scheduling and autoscaling.
  • Familiarity with LLMOps practices, model versioning, evaluation harnesses, and usage and cost observability.
  • Experience setting technical direction, leading cross-team initiatives, and mentoring engineers.
  • Clear, proactive communication skills, including the ability to explain infrastructure trade-offs to technical and non-technical stakeholders.
  • Experience operating enterprise AI-assisted developer tooling such as Claude Code or Copilot is preferred.
  • Familiarity with Okta/OIDC and enterprise authentication patterns is preferred.
  • Experience with engineering productivity metrics tools such as LinearB and AI-based code review tools such as Qodo is preferred.
  • Production experience with vector databases and RAG pipelines, including Milvus, Pinecone, or pgvector, is preferred.
  • Exposure to model fine-tuning or lightweight training pipelines such as LoRA or QLoRA is preferred.

Benefits

  • Restricted Stock Units and Employee Stock Purchase Plan.
  • Flexible time off, paid company holidays, paid sick time, gender-neutral parental leave, and grandparent leave.
  • Medical, dental, and vision coverage; 401(k) with company match; life and disability insurance; health and dependent care FSA; and voluntary insurance benefits.
  • Employee Assistance Program, prepaid legal services, pet insurance, cancer care program, and global business travel medical insurance.
  • Home office allowance and mobile phone reimbursement.
  • Wellness coach, wellness/gym reimbursement, fertility coverage, and adoption and surrogacy reimbursement.
  • SentinelOne is an Equal Employment Opportunity and Affirmative Action employer.
  • SentinelOne participates in the E-Verify Program for all U.S.-based roles.

More jobs at SentinelOne

Similar jobs