Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
AWS @ 3
Agentic AI @ 3
Communication @ 7
Data Pipelines @ 4
Databricks @ 3
Distributed Systems @ 7
GPU
LLM @ 3
MLOps
Machine Learning @ 4
Mentoring @ 4
Observability
RAG
Security
Software Development @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Stripe's ML Platform team builds platforms and services that enable machine learning engineers and data scientists to take data and build features and models from prototype to production reliably, at low latency, and at scale. The team's scope includes ML training infrastructure, model serving and deployment, feature computation and online serving, observability and monitoring, and agentic AI capabilities.
Responsibilities
- Serve as a technical lead across the ML Platform space and help define the long-term strategy and technical direction for ML infrastructure.
- Own end-to-end architecture and system design for large, complex projects across ML Platform.
- Define technical direction for ambiguous projects and translate complex user needs into long-lasting platform strategy.
- Design architectures for AI and ML workflow orchestration, scalable CPU and GPU compute infrastructure, model training, LLM fine-tuning, low-latency model inference, large-scale feature stores, real-time monitoring, and LLM and agent orchestration.
- Lead projects from requirements through design, implementation, and production operation.
- Work with ML engineers, data scientists, and product teams to translate needs into functional requirements and scalable technical solutions.
- Balance latency, reliability, cost, and security constraints when making critical technical decisions.
- Advise senior leaders on technical considerations related to the end-to-end ML lifecycle.
- Drive cross-team initiatives that improve ML development velocity and MLOps maturity.
- Mentor and grow engineers and serve as a role model for designing, implementing, and operating software systems.
Requirements
Minimum Requirements
- 10+ years of professional software development experience or equivalent domain expertise, with a strong background in service-oriented architecture and large-scale distributed systems.
- Experience serving as a technical lead, providing technical direction, leading multi-team initiatives, and mentoring team members.
- Experience building and operating production ML platforms in areas such as model training, model serving, orchestration, or ML data systems, with requirements for performance, reliability, scalability, and cost efficiency.
- Strong product instincts and understanding of the business context in which you operate.
- Strong communication skills and the ability to explain complex technical concepts to technical and non-technical stakeholders.
- Demonstrated cross-functional collaboration with ML engineers, data scientists, software engineers, product managers, and business stakeholders.
- Ability to work autonomously and responsibly in ambiguous environments.
- Hands-on experience using AI tools to accelerate work.
Preferred Qualifications
- Experience building large-scale ML training, serving, or data infrastructure, including distributed training, model inference, feature stores, real-time feature computation, and model registries.
- Experience with distributed ML training systems, accelerator-backed compute, training data pipelines, experiment tracking, and model evaluation.
- Experience rapidly developing prototypes and iterating based on user feedback.
- Experience training and shipping machine learning models to production for critical business problems.
- Familiarity with LLMs, LLM application frameworks, and agentic AI patterns such as tool use, multi-agent orchestration, and retrieval-augmented generation.
- Familiarity with AWS and cloud-based AI and ML services such as SageMaker, Bedrock, Databricks, and OpenAI.
- Ability to synthesize ideas across an organization while setting a compelling technical vision.
- Comfort working with geographically distributed teams.
- Passion for side projects, open source, or self-driven technical initiatives.
More jobs at Stripe
Full Stack Engineer, Developer & End User Experience Platform
Stripe · Toronto, Canada
CAD 135,200-202,800 per year
Head of Enterprise Solutions Architecture Platforms
Stripe · South San Francisco, United States
USD 299,800-449,600 per year
Technical Solutions Engineer
Stripe · United States
USD 134,600-201,800 per year
Strategic Programs Lead, New Markets
Stripe · Singapore, Singapore
SGD 159,200-238,800 per year
Financial Connections TechOps Manager
Stripe · Toronto, Canada
CAD 149,200-223,800 per year
Similar jobs
Staff Software Engineer, Machine Learning Platform
Stripe · Canada, Toronto, Canada
CAD 208,000-312,000 per year
Principal Engineer, AI Tooling and Workflows
Nvidia · Santa Clara, United States
USD 272,000-431,200 per year
Senior Software Engineer
SentinelOne · United States
USD 132,000-182,000 per year
Forward Deployed Engineer - Physical AI Cloud Platform
Nebius · United States, Austin, United States
USD 179,500-224,300 per year
Senior Manager, AI Software Engineering
SentinelOne · United States
USD 200,000-275,000 per year
Principal Engineer - Enterprise Content and AI Data Platform
Nvidia · Santa Clara, United States
USD 248,000-391,000 per year
NCX Senior Engineer
Nvidia · Santa Clara, United States
USD 184,000-356,500 per year
Senior Director, Enterprise Networking
Nvidia · Santa Clara, United States
USD 332,000-500,200 per year