Senior Site Reliability Engineer

at Nebius
USD 147,200-224,000 per year
SENIOR
✅ On-site

Tech Stack

AI API @ 4 CI/CD Data Pipelines DevOps @ 7 Distributed Systems @ 4 IaC Kubernetes @ 7 Networking @ 4 Observability @ 4 RAG SRE Terraform @ 6

Details

Nebius is building a full-stack AI cloud platform supporting developers and enterprises from data and model training through production deployment. Tavily is building the infrastructure layer for agentic web interaction at scale, powering Retrieval-Augmented Generation and real-time reasoning in AI systems.

This role requires working in the office in New York City and involves owning the infrastructure supporting AI agents, large-scale web workloads, and millions of daily requests.

Responsibilities

  • Manage Kubernetes clusters across multiple environments and regions.
  • Own infrastructure as code for all resources.
  • Maintain and improve CI/CD pipelines and GitOps-based deployments.
  • Maintain and optimize real-time data pipelines processing billions of events per day across distributed queues and stream processors.
  • Build monitoring, alerting, and observability systems.
  • Debug production issues across services.
  • Manage cloud costs and capacity planning.
  • Work closely with a small engineering team while owning the entire infrastructure.

Requirements

  • 5–8 years of experience in a DevOps or Site Reliability Engineering role, working in production environments.
  • Proven experience designing and operating large-scale distributed systems, with a solid understanding of API design, reliability, and performance at scale.
  • Strong Kubernetes experience in a managed cloud environment.
  • Proficiency with infrastructure as code, such as Terraform.
  • Experience with GitOps-based deployment workflows.
  • Experience building or maintaining observability stacks for logging, metrics, and alerting.
  • Experience handling production incidents calmly and methodically.

Nice to Have

  • Experience with multi-region deployments.
  • Search infrastructure experience.
  • Data pipeline experience, including streaming and warehousing.
  • Proxy or networking infrastructure experience at scale.

Benefits

  • Base compensation range of $147,200–$224,000 USD per year.
  • 100% company-paid medical, dental, and vision coverage for employees and families.
  • 401(k) plan with up to 4% company match and immediate vesting.
  • 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
  • Up to $85 per month reimbursement for mobile and internet expenses.
  • Company-paid short-term disability, long-term disability, and life insurance coverage.
  • Career growth and learning opportunities.
  • Flexibility and ownership.
  • Collaborative and innovative culture.
  • Opportunity to work on impactful AI projects.
  • International environment and talented teams.

Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire. Nebius is an equal opportunity employer.

More jobs at Nebius

Similar jobs