Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AWS @ 4
Ansible @ 6
CI/CD @ 4
Compliance @ 4
Distributed Systems @ 4
GCP @ 4
GitHub @ 4
GitHub Actions @ 4
Go @ 7
Grafana @ 7
IaC
Java @ 7
Kubernetes @ 4
Observability @ 7
OpenTelemetry @ 7
Prometheus @ 7
Python @ 7
SRE
Security
Terraform @ 6
Thanos @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
As a Staff Infrastructure Engineer, you will be a pivotal technical leader and architect within SentinelOne's Observability team. You will design, implement, and optimize critical infrastructure supporting the company's global platform, with a focus on high-volume telemetry ingestion, storage, analysis, and real-time platform visibility. The role includes end-to-end ownership of observability infrastructure, technical mentorship, cross-functional collaboration, and incident response.
U.S. citizenship is required due to Federal Government contract requirements. FedRAMP staff may be subject to customer or third-party background checks up to and including Secret Clearance if required by the role.
Responsibilities
- Architect and implement robust, scalable telemetry platforms that enable engineers to deploy and monitor features with speed, safety, and reliability.
- Serve as the primary subject matter expert and administrator for the observability stack, including Grafana, Prometheus, Thanos/Mimir/Cortex, and OpenTelemetry (OTEL) pipelines.
- Partner with engineering teams to define platform requirements and evolve the observability ecosystem ahead of stakeholder needs.
- Own critical features from architectural design and requirements refinement through production deployment and operational maturity.
- Operate observability services across AWS and GCP while balancing system reliability with cloud cost optimization.
- Build automation and self-service tooling to reduce operational toil, optimize resource utilization, and minimize pager fatigue.
- Deploy, maintain, and ensure compliance of observability systems in high-security environments, including FedRAMP and air-gapped deployments.
- Implement infrastructure as code using Terraform and Ansible and standardize industry best practices.
- Mentor engineers, lead technical design and code reviews, and provide constructive feedback.
- Resolve complex production incidents, conduct root-cause analyses, and participate in on-call rotations.
Requirements
- 8+ years of experience in infrastructure engineering, site reliability engineering, or a related systems-focused field.
- 8+ years of experience architecting, scaling, and managing enterprise-grade observability stacks using Prometheus, Grafana, Thanos, Mimir or Cortex, and OpenTelemetry.
- Experience designing cloud-native infrastructure on AWS or GCP.
- Experience managing production Kubernetes environments, including EKS and GKE.
- Advanced proficiency with Terraform and Ansible for managing immutable infrastructure.
- Experience maintaining and optimizing high-throughput, large-scale distributed systems, with a focus on cost efficiency, scalability, and disaster recovery.
- Ability to lead complex technical designs, mentor engineers, and collaborate with product and application teams.
- U.S. citizenship and the ability to work in a government-regulated environment.
Preferred Qualifications
- 8+ years of production-level programming experience in Go or another mainstream language such as Python or Java, with a willingness to adopt Go.
- Experience with FedRAMP or other sovereign cloud compliance requirements.
- Familiarity with on-premises, hybrid, or air-gapped Kubernetes deployments.
- Experience designing CI/CD pipelines, such as GitHub Actions, and implementing canary, blue-green, and rolling deployment strategies.
Benefits
- Restricted Stock Units (RSUs)
- Employee Stock Purchase Plan (ESPP)
- Flexible time off, paid company holidays, and paid sick time
- Gender-neutral parental leave and grandparent leave
- Medical, dental, and vision coverage
- 401(k) retirement plan with company match
- Life and disability insurance
- Health and dependent care FSA
- Voluntary benefits, including hospital, accident, and critical illness coverage
- Employee Assistance Program (EAP)
- ARAG prepaid legal services
- Nationwide pet insurance
- Cancer Care program
- Global business travel medical insurance
- Home office allowance and mobile phone reimbursement
- Wellness coach and wellness/gym reimbursement
- Fertility coverage and adoption/surrogacy reimbursement
The U.S. base pay range varies by candidate location. For some locations, a different pay range may apply and will be provided during the recruiting process.
Base Salary Range: $132,000–$215,000 USD