Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AWS @ 4
Ansible @ 6
CI/CD @ 4
Compliance @ 4
Distributed Systems @ 4
GCP @ 4
GitHub @ 4
GitHub Actions @ 4
Go @ 7
Grafana @ 7
IaC @ 6
Java @ 7
Kubernetes @ 4
Mentoring
Observability @ 7
OpenTelemetry @ 7
Prometheus @ 7
Python @ 7
SRE @ 7
Security @ 4
Swift
Terraform @ 6
Thanos @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Responsibilities
- Architect and implement robust, scalable telemetry platforms that empower SentinelOne engineers to deploy and monitor features with speed, safety, and reliability.
- Act as the primary Subject Matter Expert (SME) and administrator for core observability stack, including Grafana, Prometheus, Thanos/Mimir/Cortex, and OpenTelemetry (OTEL) pipelines.
- Partner strategically with diverse engineering teams across the organization to define platform requirements and evolve the observability ecosystem ahead of stakeholder needs.
- Take complete ownership of critical features, from initial architectural design and requirements refinement through production deployment and operational maturity.
- Drive operational efficiency for critical observability services across AWS and GCP, balancing system reliability with cloud cost-optimization.
- Build robust automation and self-service tooling to reduce operational toil, optimize resource utilization, and minimize pager fatigue.
- Drive deployment, maintenance, and compliance of observability systems in critical, high-security environments, including FedRAMP and air-gapped deployments.
- Cultivate platform transparency and reliability by implementing Infrastructure as Code (IaC) (Terraform/Ansible) and standardizing industry best practices.
- Elevate engineering quality by mentoring engineers, leading comprehensive technical design and code reviews, and providing constructive feedback.
- Lead swift resolution of highly complex production incidents, perform thorough root-cause analyses, and participate in on-call rotations.
Requirements
- 8+ years experience in Infrastructure Engineering, Site Reliability Engineering (SRE), or a related systems-focused field.
- 8+ years experience architecting, scaling, and managing enterprise-grade observability stacks using Prometheus, Grafana, Thanos (or Mimir/Cortex), and OpenTelemetry (OTEL).
- Experience design-engineering cloud-native infrastructure within major cloud providers (AWS or GCP) and managing production Kubernetes environments (EKS, GKE).
- Advanced proficiency with IaC and automation tools, specifically Terraform and Ansible, to manage immutable infrastructure.
- Experience maintaining and optimizing high-throughput, large-scale distributed systems with a focus on cost-efficiency, scalability, and disaster recovery.
- Demonstrated ability to lead complex technical designs, mentor other engineers, and collaborate cross-functionally with product and application teams.
- US Citizenship and ability to work in a government-regulated environment.
Preferred Qualifications
- 8+ years production-level programming experience in Go (highly desirable) or another mainstream language (e.g., Python, Java) with willingness to adopt Go.
- Experience working with high-security compliance frameworks, specifically FedRAMP or other sovereign cloud requirements.
- Familiarity with operational challenges of on-premises, hybrid, or air-gapped Kubernetes deployments.
- Experience designing advanced CI/CD pipelines (e.g., GitHub Actions) and implementing sophisticated deployment strategies (canary, blue-green, rolling updates).
Benefits
Equity & Rewards
- Restricted Stock Units (RSUs)
- Employee Stock Purchase Plan (ESPP)
Time Off & Wellbeing
- Flexible time off
- Paid company holidays and paid sick time
- Gender-neutral parental leave
- Grandparent leave
Insurance & Financial Security
- Medical, dental, and vision coverage
- 401(k) retirement plan with company match
- Life and disability insurance
- Health and dependent care FSA
- Voluntary benefits (hospital, accident, critical illness)
- Employee Assistance Program (EAP)
- ARAG pre-paid legal
- Nationwide pet insurance
- Cancer Care program
- Global business travel medical insurance
Work Perks & Flexibility
- Home office allowance
- Mobile phone reimbursement
Wellness & Lifestyle
- Wellness coach
- Wellness/gym reimbursement
- Fertility coverage
- Adoption & surrogacy reimbursement
More jobs at SentinelOne
Professional Services Project Manager
SentinelOne · United States
USD 132,000-165,000 per year
Senior Staff Ai Platform Engineer
SentinelOne · United States
USD 184,000-253,000 per year
Senior IT Product Manager, AI Initiatives
SentinelOne · United States
USD 156,000-215,000 per year
Manager Of Developer Tooling
SentinelOne · United States
USD 164,000-226,000 per year
Staff Forward Deployed Engineer, AI
SentinelOne · United States
USD 156,000-215,000 per year
Similar jobs
Senior Site Reliability Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Senior SRE Engineer
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Senior Storage Production Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 176,000-333,500 per year
Staff Software Engineer - Databases SRE
Grafana Labs · Sweden
SEK 878,600-1,054,300 per year
Senior Systems Engineer, Storage - DGX Cloud
Nvidia · United States
USD 208,000-414,000 per year
Principal Software Engineer - DGX Cloud
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year