Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
AWS @ 4
Apigee @ 4
ArgoCD @ 7
CI/CD @ 7
Claude Code @ 4
Communication @ 6
DevOps @ 7
GPU @ 4
GitHub @ 4
GitHub Actions
Jenkins @ 7
Kubernetes @ 7
LLM @ 4
LLMOps @ 3
Mentoring @ 4
Observability @ 4
RAG @ 4
Security @ 4
Terraform @ 4
Vector Databases @ 4
vLLM
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
SentinelOne is seeking a Staff AI Platform Engineer, Infrastructure Services, to own its AI Gateway infrastructure, built on Kong AI Gateway. The platform authenticates, routes, rate-limits, and monitors AI coding assistant traffic across the organization. This high-autonomy role will set technical direction for AI infrastructure, lead incident response and reliability efforts, and design solutions spanning CI/CD, GitOps, Kubernetes, and artifact systems.
Responsibilities
- Architect, harden, and scale the Kong AI Gateway deployment using Konnect Hybrid on KCP/EKS.
- Manage authentication with Okta/OIDC, consumer tiers and budgets, rate limiting, semantic caching, and observability.
- Lead root-cause analysis and remediation for gateway issues involving timeouts, latency, capacity, and failover.
- Build monitoring and alerting to identify issues before they affect users.
- Design solutions across the broader platform, including Jenkins, JPAAS, ArgoCD, Kubernetes, Artifactory, Xray, GitHub Enterprise, and GitHub Actions runner fleets.
- Evaluate and deploy AI developer tooling, including AI-assisted pull request review tools such as Qodo and engineering metrics platforms such as LinearB.
- Make build-versus-buy recommendations for AI developer tooling.
- Define architecture and standards for AI infrastructure, review team designs, and mentor engineers.
- Partner with security, Developer Experience, and product engineering teams to translate platform requirements into capabilities.
- Deploy and operate self-hosted and open-weight model serving infrastructure using technologies such as vLLM, NVIDIA Triton/NIM, TGI, and Ollama.
- Plan GPU capacity, configure autoscaling, and tune cost and performance.
- Help establish LLMOps practices covering model versioning, evaluation, safe rollout, retrieval-augmented generation, vector stores, and embedding pipelines.
- Build observability for token usage, latency, and spending across API-based and self-hosted models.
Requirements
- At least 8 years of experience in platform, infrastructure, or DevOps engineering, with a history of owning production systems end to end.
- Hands-on experience with API gateway technologies such as Kong, Envoy, Apigee, or similar.
- Experience with AI or LLM gateway patterns, including rate limiting, semantic caching, and prompt/response observability, is preferred.
- Strong Kubernetes and GitOps experience, including ArgoCD or a comparable platform.
- Experience operating across development, government, and production environments.
- Strong CI/CD experience, including Jenkins pipeline design and administration, build infrastructure, and runner or agent fleet management.
- Experience with artifact and package management systems such as Artifactory and Xray, or similar tools.
- Experience administering source control platforms such as GitHub Enterprise.
- Working knowledge of Terraform and cloud platforms such as AWS/EKS.
- Experience deploying and operating self-hosted LLM inference stacks and GPU-backed infrastructure, including Kubernetes GPU scheduling and autoscaling.
- Familiarity with LLMOps practices, including model versioning, evaluation harnesses, and usage and cost observability.
- Experience setting technical direction, leading cross-team initiatives, and mentoring engineers.
- Clear and proactive communication skills, with the ability to explain infrastructure trade-offs to technical and non-technical stakeholders.
- Enterprise experience operating LLM- or AI-assisted developer tools such as Claude Code or Copilot is preferred.
- Familiarity with Okta/OIDC and enterprise authentication patterns is preferred.
- Experience with engineering productivity metrics tools such as LinearB and AI-based code review tools such as Qodo is preferred.
- Production experience with vector databases and RAG pipelines, such as Milvus, Pinecone, or pgvector, is preferred.
- Exposure to model fine-tuning or lightweight training pipelines, such as LoRA or QLoRA, is preferred.
Benefits
- Restricted Stock Units and Employee Stock Purchase Plan.
- Flexible time off, paid company holidays, paid sick time, gender-neutral parental leave, and grandparent leave.
- Medical, dental, and vision coverage.
- 401(k) retirement plan with company match.
- Life and disability insurance.
- Health and dependent care FSA.
- Voluntary hospital, accident, and critical illness benefits.
- Employee Assistance Program and ARAG pre-paid legal services.
- Nationwide pet insurance and Cancer Care program.
- Global business travel medical insurance.
- Home office allowance and mobile phone reimbursement.
- Wellness coach, wellness/gym reimbursement, fertility coverage, and adoption and surrogacy reimbursement.
Compensation
The U.S. base salary range is $156,000–$215,000 per year and may vary based on the candidate's location. Different pay ranges may apply in some locations.
More jobs at SentinelOne
Senior AI Platform Engineer, Infrastructure Services
SentinelOne · United States
USD 132,000-182,000 per year
Manager, Detection Engineering
SentinelOne · United States
USD 164,000-226,000 per year
Staff Solutions Engineer - SLED
SentinelOne · United States
USD 224,000-308,000 per year
Staff Software Engineer, Endpoint Escalations
SentinelOne · United States
USD 156,000-215,000 per year
Staff Backend Software Engineer, Agent Platform
SentinelOne · United States
USD 156,000-215,000 per year
Similar jobs
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Infrastructure Software Engineer, TensorRT Edge-LLM
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior Data Architect
Eneco · Rotterdam, Netherlands
EUR 93,000-150,000 per year
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Forward Deployment Engineering Manager
Nebius · United States
USD 225,800-281,000 per year
Manager of Developer Tooling
SentinelOne · United States
USD 164,000-226,000 per year
Senior SOCD Applied AI Engineer
Nvidia · Santa Clara, United States
USD 168,000-310,500 per year
Senior AI Engineer
Grafana Labs · United States
USD 154,400-185,300 per year