Senior Site Reliability Engineer - US

USD 222,000-326,000 per year
SENIOR
✅ Remote

Tech Stack

AI @ 4 AWS @ 4 Communication @ 6 DevOps @ 6 GCP @ 4 GitHub Go @ 4 Grafana @ 4 Kubernetes @ 4 Linux @ 7 Networking @ 7 Observability @ 4 Prometheus @ 4 SRE @ 6 Security @ 7 Slack

Details

Teleport is the AI Infrastructure Identity Company, building secure identity solutions for humans, machines, workloads, and AI agents. The company is remote-first and globally distributed, working with organizations including Nasdaq, IBM, and Elastic.

Teleport Cloud provides a SaaS option for Teleport's traditionally open-source and enterprise access plane. The team is building production and SaaS infrastructure from scratch, solving challenging problems around security, reliability, global scale, and engineering productivity. Most of the code written for this role will be in Go.

Responsibilities

  • Re-engineer the core Teleport product to scale globally and optimize routing latency for teams distributed around the world.
  • Rewrite portions of the core Teleport product to support cloud product goals.
  • Build the monitoring and observability stack to detect production issues and minimize false positives.
  • Develop automation to eliminate high-toil activities.
  • Handle operational challenges including patching, scaling, backup and restore, and disaster recovery.
  • Investigate outages and incidents experienced by customers.
  • Participate in the on-call rotation to support 24/7/365 system uptime.
  • Operate and support the observability platform to maintain visibility and reliability.

Requirements

  • 5+ years of progressive experience in software engineering and/or SRE/DevOps roles.
  • Strong experience with Linux systems, networking, containers, and troubleshooting.
  • Solid Go and Kubernetes development experience.
  • Experience developing scripts, automation, or lightweight programs; submitting patches to product codebases; or building tooling that incorporates AI agents into operational workflows.
  • AWS Cloud experience preferred; GCP experience is acceptable.
  • Experience with systems observability tools such as Prometheus, Grafana, and Loki.
  • Experience working in environments where strong security choices are critical and where correctness and system invariants, including formal or property-based methods, are valued.
  • Willingness to collaborate with Teleport engineers on a Go coding challenge during the interview process.
  • Intellectual curiosity and willingness to master new technologies.
  • Transparency, honesty, and a no-ego mindset.
  • Excellent communication skills.

Interview Process

  1. Zoom meeting with a Teleport recruiter covering the company, products, compensation philosophy, interview process, and key requirements.
  2. Zoom meeting with the hiring manager or a lead engineer to review the coding challenge.
  3. Collaboration on a Go coding challenge using GitHub. Candidates join a Slack channel and typically have about two weeks to complete the challenge, followed by a Zoom review with the hiring manager and an interview team member.
  4. Candidates whose solution meets Teleport's requirements receive an offer.

Important Information

  • The role is remote in the United States, but requires in-person attendance at the Oakland, California office during onboarding week.
  • Background checks are conducted.

Benefits

  • Extensive health coverage.
  • Annual expense budget.
  • Rest and recovery policies.
  • Retirement savings plans.
  • Professional development opportunities.
  • Equity offered as part of compensation.

Teleport is an equal opportunity employer and does not discriminate against employees or applicants on the basis of protected characteristics.

More jobs at Teleport

Similar jobs