Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Distributed Systems @ 6
Networking @ 6
Observability @ 4
SRE
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Reddit's Ads Reliability team partners with Ads Engineering to improve reliability, scalability, operational excellence, and developer productivity across the advertising ecosystem. The team builds and operates highly available services supporting Reddit Ads, including ad-serving, auction, targeting, measurement, and billing systems.
Responsibilities
- Partner with Ads Engineering teams to improve the reliability, scalability, and operational excellence of ad-serving, auction, targeting, measurement, and billing systems.
- Design, build, and maintain infrastructure, tooling, and automation that improve service reliability and engineering productivity.
- Improve observability through monitoring, alerting, tracing, logging, and dashboards.
- Participate in on-call rotations and lead incident response efforts for critical production systems.
- Perform root cause analysis and drive corrective actions following incidents.
- Collaborate with software engineers throughout the service lifecycle, from design reviews through production operations.
- Drive adoption of SRE best practices, including SLIs, SLOs, error budgets, capacity planning, and operational readiness reviews.
- Reduce operational toil through automation and self-service tooling.
- Help define and measure advertiser-critical user journeys, such as campaign creation, ad delivery, reporting, and billing.
- Scale Ads systems to support continued traffic growth, increased advertiser demand, and evolving business requirements.
Requirements
- 5+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or related roles operating large-scale distributed systems.
- Strong experience supporting high-traffic, user-facing production environments.
- Strong cross-functional collaboration skills with the ability to lead and influence operational excellence.
- Good understanding of modern distributed systems, scale engineering, and cloud-native architectures.
- Strong software engineering skills in general-purpose backend languages such as Go.
- Demonstrated ability to troubleshoot complex issues across applications, infrastructure, networking, and services.
- Experience with observability platforms, monitoring systems, alerting, and incident response.
- Experience driving automation and operational improvements.
Benefits
- Comprehensive health benefits
- 401(k) matching
- Workspace benefits for a home office
- Personal and professional development funds
- Family planning support
- Flexible vacation and Reddit Global Days Off
- 4+ months of paid parental leave
- Paid volunteer time off
The position is eligible for equity in the form of restricted stock units and may also be eligible for a commission depending on the position offered. U.S.-based benefits include medical, dental, and vision insurance, a 401(k) program with employer match, vacation time, and parental leave.
More jobs at Reddit
Senior Frontend Engineer, Ads Creative
Reddit · United States
USD 190,800-267,100 per year
Staff Software Engineer - Site Defense
Reddit · United States
USD 217,000-303,900 per year
Product Manager, Developer Ecosystem
Reddit · United States
USD 217,000-303,900 per year
Staff Product Manager, Ads Formats
Reddit · United States, Chicago, United States, New York City, United States, Los Angeles, United States, San Francisco, United States
USD 217,000-303,900 per year
Engineering Manager, Notifications Platform
Reddit · United States
USD 217,000-303,000 per year
Similar jobs
Staff Software Engineer, AI Reliability
Anthropic · New York City, United States, San Francisco, United States, Seattle, United States
USD 325,000-485,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity AI · Palo Alto, United States, San Francisco, United States
USD 220,000-405,000 per year
Senior Engineering Manager, Infrastructure Security Engineering - DGX Cloud
Nvidia · United States
USD 248,000-391,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Staff Forward Deployed Engineer, Agentic SDLC
GitLab · United States
USD 254,000-297,000 per year
Senior Site Reliability Engineer
Nebius · New York City, United States
USD 147,200-224,000 per year
Senior Software Engineer - Public Cloud Engineering
Bloomberg · New York City, United States
USD 160,000-240,000 per year