Member of Technical Staff (Software Engineer, Data Platform)
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API
Airflow @ 3
ClickHouse
Dagster @ 3
Data Modeling @ 4
Data Pipelines
Data Science
Databricks
Experimentation
Flink
Go @ 6
Kafka
Kinesis
Machine Learning
Observability @ 3
Python @ 6
Snowflake
Spark
Streaming Data Processing @ 4
TypeScript @ 6
dbt
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
The Data Platform team owns the end-to-end data lifecycle at Perplexity, from ingestion through processing, storage, and serving. The platform powers product features, analytics, experimentation, AI workloads, and the company’s data lake.
The team defines the architecture for batch and streaming systems, orchestration and observability, and a self-service data platform. It combines Databricks and Snowflake with open-source technologies including Spark, Kafka, Flink, Airflow, Dagster, dbt, Iceberg, Delta Lake, and ClickHouse. In this senior/staff role, you will shape architecture, set standards, and drive the long-term technical direction of Perplexity’s data ecosystem.
Responsibilities
- Design and operate large-scale batch and streaming data pipelines that power product features, AI training and evaluation workflows, analytics, and experimentation.
- Build event-driven and streaming systems using Kafka, Kinesis, Pub/Sub, or similar technologies for real-time ingestion, transformation, and delivery.
- Build batch frameworks for backfills, aggregations, and offline computation.
- Lead data orchestration architecture using Airflow, Dagster, or similar tools, including scheduling, dependency management, retries, SLAs, and end-to-end observability.
- Set and enforce guarantees for data correctness, freshness, lineage, and recoverability.
- Design systems that handle rapid scale growth, partial failures, and evolving schemas without disrupting AI workloads or product experiences.
- Build self-service data platforms that enable engineers, data scientists, and analysts to discover data, define contracts, and create and operate pipelines.
- Improve developer experience through abstractions, opinionated paved paths, and standards for data modeling, testing, validation, and deployment.
- Drive architectural decisions across storage, compute, orchestration, and data APIs.
- Partner with product engineering and data science to align the data ecosystem with the company roadmap.
- Mentor engineers, review designs, and raise the technical bar for data infrastructure through feedback, documentation, and hands-on collaboration.
Requirements
- 5+ years of software engineering experience for Senior level or 8+ years for Staff level.
- Strong experience building production data infrastructure systems.
- Hands-on experience with batch and/or streaming data processing at scale.
- Deep familiarity with data orchestration systems such as Airflow or Dagster.
- Proficiency in Python and at least one additional backend language, such as Go or TypeScript.
- Strong systems thinking around reliability, latency, cost, and complexity tradeoffs.
- Experience supporting ML/AI workflows, training pipelines, or evaluation systems.
- Familiarity with data quality, lineage, observability, and governance tooling.
- Prior ownership of internal platforms used by many teams.
Applicants are encouraged to apply even if their experience does not match every listed qualification.
Benefits
- Equity.
- U.S. full-time employees receive health, dental, vision, retirement, fitness, commuter, and dependent care benefits, among others.
- International full-time employees receive benefits tailored to their region of residence.
- USD salary ranges apply only to U.S.-based positions. International salaries are set based on the local market. Final offer amounts depend on factors including experience and expertise.