Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AWS @ 4
Bash @ 6
Compliance @ 4
DevOps @ 7
GCP @ 4
Go @ 6
Grafana @ 4
Kubernetes @ 4
Observability @ 4
OpenTelemetry @ 4
Prometheus @ 4
Python @ 6
Ruby @ 6
SRE @ 7
Security @ 4
Terraform @ 3
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
As a Staff Site Reliability Engineer, you will join the Government SRE team and own the technical reliability of government environments, as well as the coordination of compliant and efficient deployments. You will work at the intersection of regulated-environment compliance requirements and the need to establish a consistent software experience for users and developers in commercial environments. The role involves close collaboration with security, compliance, operations, validation, and engineering teams to lead best practices for cloud infrastructure, continuous delivery, and government release processes.
Due to Federal Government contract requirements, U.S. Citizenship and a work location in the United States are required. FedRAMP staff may be subject to customer or third-party background checks up to and including Secret Clearance if required by their role at SentinelOne.
Responsibilities
- Drive continuous software delivery, resolve incidents, run postmortems, and create automation strategies for deployment, self-testing, and alerting.
- Lead and execute incident management for production issues, ensuring rapid recovery, root cause analysis, and preventative follow-up actions.
- Improve and optimize observability strategies by collaborating with application engineering teams to design monitoring solutions that enhance alerting capabilities and reduce noise.
- Define, implement, and monitor SLOs, SLIs, and SLAs in collaboration with product and engineering teams to align with business objectives.
- Design, develop, and maintain software solutions addressing operational, compliance, and pipeline challenges.
- Own and coordinate government environment releases, driving process improvements to enhance release pipeline efficiency, reliability, and visibility.
- Understand product architecture and service dependencies to manage risk and implement effective testing strategies.
- Partner with engineering, product, SecOps, compliance, and leadership teams to align priorities, define testing strategies, and resolve challenges.
- Ensure infrastructure and deployments meet FedRAMP, government regulations, and industry standards while maintaining required release documentation and risk assessments.
Requirements
- 8+ years of experience in SRE, DevOps, or infrastructure engineering for SaaS products, including 4+ years running operations at large scale.
- 2+ years of production experience with a container orchestration system, preferably Kubernetes, and continuous delivery.
- Strong understanding of compliance frameworks relevant to government deployments, including FedRAMP, DoD, NIST 800-53, and NIST 800-137.
- Multi-cloud experience with AWS and GCP; AWS expertise is preferred.
- Experience with at least one primary programming language, such as Python, Go, or Ruby, and proficiency in Bash scripting.
- Familiarity with GitOps frameworks, infrastructure-as-code tooling such as Terraform or Pulumi, and deployment strategies including blue-green, rolling, and canary deployments.
- Experience with industry-standard observability stacks such as Prometheus, Grafana, ELK, and OpenTelemetry.
- Experience with incident management processes.
- Proven experience implementing and supporting FedRAMP, security, risk management, and compliance processes for software releases.
- Experience working directly with government agencies or in highly regulated industries.
- Familiarity with testing strategies and automation in large-scale environments.
Benefits
- Restricted Stock Units (RSUs) and Employee Stock Purchase Plan (ESPP).
- Flexible time off, paid company holidays, paid sick time, gender-neutral parental leave, and grandparent leave.
- Medical, dental, and vision coverage.
- 401(k) retirement plan with company match.
- Life and disability insurance.
- Health and dependent care FSA.
- Voluntary hospital, accident, and critical illness benefits.
- Employee Assistance Program and ARAG prepaid legal services.
- Nationwide pet insurance, Cancer Care program, and global business travel medical insurance.
- Home office allowance and mobile phone reimbursement.
- Wellness coach, wellness/gym reimbursement, fertility coverage, and adoption and surrogacy reimbursement.
Compensation
The U.S. base salary range is $156,000–$200,000 per year. The applicable range may vary based on the candidate's location; a different range may be provided during the recruiting process.