Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AWS @ 4
ArgoCD
Cassandra @ 4
ClickHouse @ 4
Communication @ 6
Distributed Systems @ 4
Docker @ 4
GCP @ 4
GitHub
Go @ 6
GraphQL
Helm @ 4
Java @ 6
Kafka @ 4
Kubernetes @ 4
Linux
Observability
PostgreSQL @ 4
Python @ 6
Redis @ 4
Security
gRPC
macOS
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
As a Senior Software Engineer on the Agent Platform team, you will help secure tens of millions of devices across Windows, Linux, and macOS while processing billions of security events every day as part of SentinelOne’s Endpoint Protection product line.
The role focuses on solving complex distributed systems challenges in production, responding to customer-critical incidents, and improving the reliability, operability, scalability, and observability of the Agent Platform. You will build and evolve high-throughput, highly available services responsible for policy, configuration, and command distribution to millions of agents worldwide.
Responsibilities
- Drive rapid response to customer-critical incidents by diagnosing, triaging, and resolving complex production issues spanning multiple Agent Platform services.
- Lead systematic root-cause analysis and translate incident learnings into durable reliability, scalability, observability, and operability improvements.
- Build a systems-level understanding of unfamiliar services and codebases while collaborating with engineering teams to resolve cross-service issues.
- Design, develop, test, document, deploy, and operate large-scale, high-volume, low-latency distributed systems processing millions of events per second.
- Maintain and improve existing services through refactoring, feature development, and architectural enhancements.
- Monitor key metrics, improve operational visibility, and strengthen application stability and data integrity.
- Translate business and functional requirements into robust, scalable, and operable technical solutions.
- Partner with engineering teams across SentinelOne to solve cross-functional problems, influence technical direction, and deliver scalable solutions.
- Evaluate and adopt technologies that improve platform scalability, reliability, and operational excellence.
- Contribute to agent platform protocols that enable SentinelOne teams to deliver new security capabilities safely and at scale.
Requirements
- 5+ years of professional backend software engineering experience.
- Deep expertise in at least one of Java, Go, or Python, with willingness to work across all three.
- Reliability-first mindset and proven ability to solve complex production incidents, perform systematic root-cause analysis, and implement durable solutions.
- Experience understanding unfamiliar distributed systems, navigating large codebases, and troubleshooting failures across service boundaries.
- Strong hands-on experience designing, building, and operating large-scale distributed systems.
- Deep understanding of failure modes, performance trade-offs, resilience patterns, and operational excellence.
- Experience with AWS, GCP, or similar cloud platforms, as well as Docker, Helm, and Kubernetes.
- Experience with messaging systems and data platforms such as Kafka, PostgreSQL, Redis, Cassandra, and ClickHouse, or similar technologies.
- Excellent communication and collaboration skills, including the ability to work across engineering teams, Product, Technical Account Managers, and other stakeholders.
- Ability to influence technical direction and mentor fellow engineers.
- High degree of ownership, autonomy, curiosity, and ability to drive ambiguous technical problems to successful outcomes.
- Experience in an enterprise SaaS or cybersecurity software company is highly desirable.
Technologies
- Languages: Java, Go, Python
- Service communication: gRPC, REST, GraphQL, Kafka
- Data platforms: Redis, PostgreSQL, Cassandra, ClickHouse, and a columnar time-series database
- Cloud and infrastructure: AWS, GCP, Kubernetes, Docker, Helm, GitHub, ArgoCD
- Supported operating systems: Windows, Linux, macOS
Benefits
- Restricted Stock Units (RSUs)
- Employee Stock Purchase Plan (ESPP)
- Flexible time off, paid company holidays, and paid sick time
- Gender-neutral parental leave and grandparent leave
- Medical, dental, and vision coverage
- 401(k) retirement plan with company match
- Life and disability insurance
- Health and dependent care FSA
- Voluntary benefits, Employee Assistance Program, pre-paid legal, pet insurance, Cancer Care program, and global business travel medical insurance
- Home office allowance and mobile phone reimbursement
- Wellness coach, wellness/gym reimbursement, fertility coverage, and adoption and surrogacy reimbursement
Compensation
The U.S. base salary range is $132,000–$182,000 USD, with the applicable range varying based on the candidate’s location. A different pay range may apply for some locations.