Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 1
API @ 4
AWS @ 4
Apigee @ 4
ArgoCD @ 7
CI/CD @ 4
Claude Code @ 4
Communication @ 6
DevOps @ 6
GPU @ 4
GitHub @ 4
GitHub Actions
Jenkins @ 4
Kubernetes @ 7
LLM @ 1
LLMOps @ 3
Mentoring @ 4
Observability @ 1
RAG @ 4
Security
Terraform @ 4
Vector Databases @ 4
vLLM
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
SentinelOne is seeking a Senior AI Platform Engineer, Infrastructure Services, to own its AI Gateway infrastructure, built on Kong AI Gateway. The platform authenticates, routes, rate-limits, and monitors AI coding assistant traffic across the organization. This high-autonomy role will set technical direction for AI infrastructure, lead incident response and reliability efforts, and design solutions spanning CI/CD, GitOps, Kubernetes, and artifact systems.
Responsibilities
- Architect, harden, and scale the Kong AI Gateway deployment using Konnect Hybrid on KCP/EKS.
- Manage authentication with Okta/OIDC, consumer tiers and budgets, rate limiting, semantic caching, and observability.
- Lead root-cause analysis and remediation for gateway issues involving timeouts, latency, capacity, and failover.
- Build monitoring and alerting to identify issues before they affect users.
- Design solutions across Jenkins, JPAAS, ArgoCD, Kubernetes, Artifactory/Xray, GitHub Enterprise, and GitHub Actions runner fleets.
- Evaluate and roll out AI developer tooling, including Qodo for AI-assisted pull request review and LinearB for engineering metrics.
- Make build-versus-buy recommendations for AI infrastructure and developer tooling.
- Define architecture and engineering standards, review designs, and mentor other engineers.
- Partner with security, DevEx, and product engineering teams to translate platform needs into capabilities.
- Deploy and operate self-hosted and open-weight model serving infrastructure using technologies such as vLLM, NVIDIA Triton/NIM, TGI, and Ollama.
- Plan GPU capacity, configure autoscaling, and tune cost and performance for model-serving workloads.
- Help establish LLMOps practices for model versioning, evaluation, and safe rollout.
- Support retrieval-augmented generation infrastructure, including vector stores and embedding pipelines.
- Build observability for token usage, latency, and spending across API-based and self-hosted models.
Requirements
- At least 5 years of experience in platform, infrastructure, or DevOps engineering, with experience owning production systems end to end.
- Hands-on experience with API gateways such as Kong, Envoy, Apigee, or similar. Experience with AI/LLM gateway patterns, including rate limiting, semantic caching, and prompt/response observability, is a strong plus.
- Strong Kubernetes and GitOps experience, including ArgoCD or a comparable platform, across development, government, and production environments.
- Solid CI/CD experience, including Jenkins pipeline design and administration, build infrastructure, and runner or agent fleet management.
- Experience with artifact and package management systems such as Artifactory and Xray, or similar tools.
- Experience administering source control platforms such as GitHub Enterprise.
- Working knowledge of Terraform and AWS/EKS.
- Experience deploying and operating self-hosted LLM inference stacks and GPU-backed infrastructure, including Kubernetes GPU scheduling and autoscaling.
- Familiarity with LLMOps practices, model versioning, evaluation harnesses, and usage and cost observability.
- Experience setting technical direction, leading cross-team initiatives, and mentoring engineers.
- Clear, proactive communication skills, including the ability to explain infrastructure trade-offs to technical and non-technical stakeholders.
- Experience operating enterprise AI-assisted developer tooling such as Claude Code or Copilot is preferred.
- Familiarity with Okta/OIDC and enterprise authentication patterns is preferred.
- Experience with engineering productivity metrics tools such as LinearB and AI-based code review tools such as Qodo is preferred.
- Production experience with vector databases and RAG pipelines, including Milvus, Pinecone, or pgvector, is preferred.
- Exposure to model fine-tuning or lightweight training pipelines such as LoRA or QLoRA is preferred.
Benefits
- Restricted Stock Units and Employee Stock Purchase Plan.
- Flexible time off, paid company holidays, paid sick time, gender-neutral parental leave, and grandparent leave.
- Medical, dental, and vision coverage; 401(k) with company match; life and disability insurance; health and dependent care FSA; and voluntary insurance benefits.
- Employee Assistance Program, prepaid legal services, pet insurance, cancer care program, and global business travel medical insurance.
- Home office allowance and mobile phone reimbursement.
- Wellness coach, wellness/gym reimbursement, fertility coverage, and adoption and surrogacy reimbursement.
- SentinelOne is an Equal Employment Opportunity and Affirmative Action employer.
- SentinelOne participates in the E-Verify Program for all U.S.-based roles.
More jobs at SentinelOne
Staff AI Platform Engineer, Infrastructure Services
SentinelOne · United States
USD 156,000-215,000 per year
Manager, Detection Engineering
SentinelOne · United States
USD 164,000-226,000 per year
Staff Solutions Engineer - SLED
SentinelOne · United States
USD 224,000-308,000 per year
Staff Software Engineer, Endpoint Escalations
SentinelOne · United States
USD 156,000-215,000 per year
Staff Backend Software Engineer, Agent Platform
SentinelOne · United States
USD 156,000-215,000 per year
Similar jobs
Senior Systems Software Engineer, Developer Productivity and Cloud Automation - GeForce NOW
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Infrastructure Software Engineer, TensorRT Edge-LLM
Nvidia · Santa Clara, United States
USD 184,000-287,500 per year
Senior Data Architect
Eneco · Rotterdam, Netherlands
EUR 93,000-150,000 per year
Senior Staff Network Automation Engineer
Nvidia · Santa Clara, United States
USD 208,000-333,500 per year
Forward Deployment Engineering Manager
Nebius · United States
USD 225,800-281,000 per year
Manager of Developer Tooling
SentinelOne · United States
USD 164,000-226,000 per year
Senior SOCD Applied AI Engineer
Nvidia · Santa Clara, United States
USD 168,000-310,500 per year
Senior AI Engineer
Grafana Labs · United States
USD 154,400-185,300 per year