Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 4
CI/CD
Data Pipelines
DevOps @ 7
Distributed Systems @ 4
IaC
Kubernetes @ 7
Networking @ 4
Observability @ 4
RAG
SRE
Terraform @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Nebius is building a full-stack AI cloud platform supporting developers and enterprises from data and model training through production deployment. Tavily is building the infrastructure layer for agentic web interaction at scale, powering Retrieval-Augmented Generation and real-time reasoning in AI systems.
This role requires working in the office in New York City and involves owning the infrastructure supporting AI agents, large-scale web workloads, and millions of daily requests.
Responsibilities
- Manage Kubernetes clusters across multiple environments and regions.
- Own infrastructure as code for all resources.
- Maintain and improve CI/CD pipelines and GitOps-based deployments.
- Maintain and optimize real-time data pipelines processing billions of events per day across distributed queues and stream processors.
- Build monitoring, alerting, and observability systems.
- Debug production issues across services.
- Manage cloud costs and capacity planning.
- Work closely with a small engineering team while owning the entire infrastructure.
Requirements
- 5–8 years of experience in a DevOps or Site Reliability Engineering role, working in production environments.
- Proven experience designing and operating large-scale distributed systems, with a solid understanding of API design, reliability, and performance at scale.
- Strong Kubernetes experience in a managed cloud environment.
- Proficiency with infrastructure as code, such as Terraform.
- Experience with GitOps-based deployment workflows.
- Experience building or maintaining observability stacks for logging, metrics, and alerting.
- Experience handling production incidents calmly and methodically.
Nice to Have
- Experience with multi-region deployments.
- Search infrastructure experience.
- Data pipeline experience, including streaming and warehousing.
- Proxy or networking infrastructure experience at scale.
Benefits
- Base compensation range of $147,200–$224,000 USD per year.
- 100% company-paid medical, dental, and vision coverage for employees and families.
- 401(k) plan with up to 4% company match and immediate vesting.
- 20 weeks of paid parental leave for primary caregivers and 12 weeks for secondary caregivers.
- Up to $85 per month reimbursement for mobile and internet expenses.
- Company-paid short-term disability, long-term disability, and life insurance coverage.
- Career growth and learning opportunities.
- Flexibility and ownership.
- Collaborative and innovative culture.
- Opportunity to work on impactful AI projects.
- International environment and talented teams.
Applicants must be authorized to work in the country in which they apply and must provide proof of employment eligibility as a condition of hire. Nebius is an equal opportunity employer.
More jobs at Nebius
Forward Deployment Engineering Manager
Nebius · United States
USD 225,800-281,000 per year
IT Risk & Controls Manager
Nebius · United States
USD 120,000-180,000 per year
IT Infrastructure Engineer – RMA & Hardware Diagnostics
Nebius · Kansas City, United States
USD 112,700-140,800 per year
Senior Machine Learning Engineer, Model Training and Reinforcement Learning
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Senior Machine Learning Engineer, LLM Inference Optimization
Nebius · Palo Alto, United States
USD 195,200-262,200 per year
Similar jobs
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Site Reliability Engineer, AIOps
Nvidia · Santa Clara, United States
USD 148,000-276,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · San Francisco, United States, Palo Alto, United States
USD 220,000-405,000 per year
Senior Manager, Storage Production Engineering
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer - Datacenter Systems
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Partner Solutions Architect
Nebius · United States, Canada
USD 250,000-320,000 per year