Staff Software Engineer, Machine Learning Platform

at Stripe
📍 Canada
📍 Toronto, Canada
CAD 208,000-312,000 per year
SENIOR
✅ Remote ✅ On-site

Tech Stack

AI @ 4 AWS @ 3 Agentic AI @ 3 Communication @ 7 Data Pipelines @ 4 Databricks @ 3 Distributed Systems @ 8 GPU LLM @ 3 MLOps Machine Learning @ 4 Mentoring @ 4 Observability RAG Security Software Development @ 8

Details

Stripe is a financial infrastructure platform for businesses. The ML Platform team builds platforms and services that enable machine learning engineers and data scientists across Stripe to take data and build features and models from prototype to production reliably, at low latency, and at scale. The team's scope includes ML training infrastructure, model serving and deployment, feature computation and online serving, observability and monitoring, and agentic AI capabilities.

Responsibilities

  • Serve as a technical lead across the ML Platform space and contribute to the evolution of platforms powering Stripe's ML-driven products.
  • Define the long-term strategy and technical direction for the next generation of ML infrastructure.
  • Own end-to-end architecture and system design for large, complex projects across ML Platform.
  • Define technical direction for ambiguous projects and transform complex user needs into long-term platform strategy.
  • Design architectures for AI and ML workflow orchestration, scalable CPU and GPU compute infrastructure, model training, LLM fine-tuning, low-latency model inference, large-scale feature stores, real-time monitoring, and LLM and agent orchestration.
  • Lead projects from requirements through design, implementation, and production operation.
  • Translate the needs of ML engineers, data scientists, and product teams into functional requirements and scalable technical solutions.
  • Balance latency, reliability, cost, and security constraints when making critical technical decisions.
  • Advise senior leaders on technical considerations related to the end-to-end ML lifecycle.
  • Drive cross-team initiatives that improve ML development velocity and MLOps maturity.
  • Mentor and grow engineers while serving as a role model for designing, implementing, and operating software systems.

Requirements

Minimum Requirements

  • 10+ years of professional software development experience or equivalent domain expertise, with a solid background in service-oriented architecture and large-scale distributed systems.
  • Experience serving as a technical lead, providing technical direction, leading multi-team initiatives, and mentoring team members.
  • Experience building and operating production ML platforms in areas such as model training, model serving, orchestration, or ML data systems, with requirements for performance, reliability, scalability, and cost efficiency.
  • Strong product instincts and understanding of the business context in which you operate.
  • Strong communication skills and the ability to explain complex technical concepts to technical and non-technical stakeholders.
  • Demonstrated ability to collaborate cross-functionally with ML engineers, data scientists, software engineers, product managers, and business stakeholders.
  • Ability to work autonomously and responsibly in ambiguous environments.
  • Hands-on experience using AI tools to accelerate work.

Preferred Qualifications

  • Experience building large-scale ML training, serving, or data infrastructure, including distributed training, model inference, feature stores, real-time feature computation, and model registries.
  • Experience with distributed ML training systems, accelerator-backed compute, training data pipelines, experiment tracking, and model evaluation.
  • Experience rapidly developing prototypes and iterating based on user feedback.
  • Experience training and shipping machine learning models to production to solve critical business problems.
  • Familiarity with LLMs, LLM application frameworks, and agentic AI patterns such as tool use, multi-agent orchestration, and retrieval-augmented generation.
  • Familiarity with AWS and cloud-based AI and ML services such as SageMaker, Bedrock, Databricks, and OpenAI.
  • Ability to synthesize ideas across an organization while setting a compelling technical vision.
  • Comfort working with geographically distributed teams.
  • Passion for side projects, open source, or self-driven technical initiatives.

More jobs at Stripe

Similar jobs