Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 4
Communication @ 4
Data Pipelines @ 7
Distributed Systems @ 8
Flink @ 4
Go @ 7
Java @ 7
Kafka @ 4
Kubernetes @ 3
Leadership @ 4
Machine Learning
Observability @ 4
Python @ 7
Scala @ 7
Security
Software Development @ 7
Technical Leadership @ 4
gRPC @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Reddit's Infrastructure organization enables reliability, performance, and efficiency across the platform through a unified technology stack. The Content Platform team within Infrastructure owns Tier-0 services and core data models supporting feeds, posting, commenting, and upvoting, as well as Reddit's R2 monolith legacy stack.
The Staff Software Engineer will lead the development and evolution of Reddit's Ingestion Platform, shaping its operational posture and incorporating feedback from platform customers and partner teams.
Responsibilities
- Design, write, and deliver reliable software for distributed data movement across streaming and batch workloads, focusing on availability, scalability, latency, correctness, and cost efficiency.
- Own the architecture and evolution of the platform's control plane and data plane, including pipeline APIs, controllers, connectors, target managers, schemas, and sink integrations.
- Expand the platform beyond Kafka-to-BigQuery by delivering production-ready paths to S3, GCS, and Apache Iceberg, along with transformations, deduplication, dead-letter queues, and other reusable capabilities.
- Build high-quality connectors and abstractions for Kafka, BigQuery, S3, GCS, Iceberg, Flink, and other data stores while keeping the platform modular and extensible.
- Improve the self-service experience through clear APIs, safe defaults, automated provisioning, documentation, onboarding workflows, actionable observability, and alerting.
- Lead migrations from bespoke and legacy ingestion systems to the Ingestion Platform, partnering with teams such as Ads Data Platform and ML Indexing.
- Establish reliability, security, and operational practices for pipelines running across Kubernetes clusters, including schema evolution, workload identity, permissions, deployment safety, monitoring, and incident response.
- Identify architectural and product experience gaps and lead redesigns that improve developer velocity and support Reddit's growth.
- Collaborate with engineers and stakeholders across Infrastructure, Data Platform, Product, Ads, ML, Storage, and partner teams.
- Mentor and guide backend and data infrastructure engineers across the company.
Requirements
- 10+ years of hands-on experience building internet-scale distributed systems, data infrastructure, or platforms used by other developers.
- BS, MS, or PhD in Computer Science or a related field, or equivalent practical experience.
- Strong software development experience in one or more general-purpose languages such as Go, Python, Java, or Scala.
- Deep experience designing and operating high-throughput, fault-tolerant data pipelines or platform services, including streaming and batch processing.
- Experience with data movement systems and patterns such as Kafka, BigQuery, S3, GCS, Apache Iceberg, Flink, or comparable technologies.
- Experience designing APIs and platform abstractions, including schema evolution, Protocol Buffers, and service-to-service communication such as gRPC.
- Familiarity with Kubernetes and cloud-native platform concepts, including controllers, custom resources, workload identity, infrastructure automation, and multi-cluster deployments.
- Strong understanding of observability and operational excellence, including metrics, logging, alerting, SLOs, incident response, data quality, and safe migration practices.
- Product-minded approach with the ability to turn ambiguous customer and partner needs into simple, reliable, and scalable solutions.
- Demonstrated technical leadership and experience guiding teams through complex, cross-functional initiatives.
- Exceptional written and verbal communication skills.
- Experience building low-code or self-service developer platforms, or curiosity to learn, is a strong plus.
Benefits
- Equity in the form of restricted stock units.
- Medical, dental, and vision insurance.
- 401(k) program with employer match.
- Generous vacation time and parental leave.
- Reasonable accommodations are available for qualified individuals with disabilities and disabled veterans.
Compensation
The base salary range is $217,000–$303,900 USD. The role may also be eligible for commission depending on the position offered. Final offers depend on skills, depth of experience, and relevant licenses or credentials.
The role is remote within the United States. In select roles and locations, interviews may be recorded, transcribed, and summarized by artificial intelligence, with an option to opt out before scheduled interviews.