Senior Manager, Platform Operations

USD 168,000-210,000 per year
SENIOR
✅ Remote

Tech Stack

AI @ 6 AWS @ 4 Ansible @ 4 Audit @ 4 Azure @ 1 Change Management @ 6 ChatGPT @ 6 Compliance @ 4 Datadog @ 4 GCP @ 4 Grafana @ 4 Kubernetes @ 4 Leadership @ 7 Observability @ 4 Prometheus @ 4 Python @ 4 Security Terraform @ 4

Details

You'll manage the team that keeps the Collibra Platform running and current for every enterprise customer around the clock, spanning incident response, change management, release execution, and fleet health. You will report to the Senior Director of Reliability Engineering & Operations and oversee the team directly, partnering closely with two technical leads across the team's two domains.

This is a highly visible, business-critical function focused on keeping customer environments healthy, current, and secure. The role supports enterprise customers and requires a US citizen residing on US soil because it supports the US government.

Responsibilities

  • Own 24x7 incident response, SaaS customer operations, and change management to deliver consistent platform reliability for enterprise customers.
  • Direct deployment and fleet operations across thousands of virtual machines underpinning the Collibra Platform, including weekly release delivery.
  • Build succession and development plans across the team's two technical domains.
  • Advance AI-powered automation across incident remediation, release delivery, and internal tooling to reduce manual toil and modernize operations.
  • Partner with engineering, product, security, finance, and support organizations to align priorities, resolve dependencies, and maintain compliance in a regulated environment.
  • Participate in incident response and release cycles as a hands-on contributor.
  • Establish operational health reporting covering MTTR, MTTD, and change failure rate.
  • Implement AI-driven improvements and shape the team's operating model for an AI-first future.

Requirements

  • 7+ years of engineering experience, including at least 3+ years in a leadership or management role overseeing incident response, release, or deployment operations for customer-facing SaaS production environments.
  • Experience managing 24x7 production operations for customer-facing systems, including on-call rotation and escalation models.
  • Experience operating within a FedRAMP or comparable regulated compliance environment, including continuous monitoring and audit cycles.
  • Experience with cloud infrastructure at enterprise scale, including AWS, AWS GovCloud, and GCP. Azure experience is a plus.
  • Experience managing distributed infrastructure fleets, virtual machines, containers, or equivalent systems supporting production SaaS environments.
  • Demonstrated proficiency using AI tools such as Claude, Gemini, ChatGPT, or Copilot to solve business challenges, drive measurable outcomes, or streamline workflows.
  • Bachelor's degree or equivalent related working experience.
  • Must be a US citizen residing on US soil.
  • Ability to balance hands-on technical depth with people leadership, operate with urgency under production pressure, maintain rollback discipline, think strategically, and communicate operational health directly and blamelessly.
  • Experience with infrastructure automation and observability tooling such as Ansible, Terraform, Python, Kubernetes, Datadog, or Prometheus/Grafana.

Measures of Success

Within the First Month

  • Build working relationships with both technical leads and the full team across the two operational domains.
  • Gain fluency in incident response, release, and change management processes, including FedRAMP audit and control obligations.
  • Review active projects, team development, succession planning, and operational health reporting.

Within the Third Month

  • Participate in incident response and release cycles as a hands-on contributor.
  • Put development and succession plans in motion for both technical domains.
  • Identify and begin scoping a workflow suited for AI-driven automation.
  • Establish regular operational health reporting for leadership visibility.
  • Build partnerships across engineering, product, security, finance, and support.

Within the Sixth Month

  • Demonstrate measurable progress against development and succession plans.
  • Implement at least one AI-driven improvement to incident remediation, release delivery, or internal tooling.
  • Improve at least one operational health metric, such as MTTR or change failure rate.
  • Document a perspective on how the team's operating model should evolve for an AI-first future.

Compensation

The standard base salary range is $168,000–$210,000 per year. The position is not eligible for additional commission-based compensation. Offers are based on factors including experience, skills, and location. Additional total rewards may include bonus potential, equity for eligible roles, a Flex Fund monthly stipend, pension or 401(k) plans, health coverage, and time off.

Benefits

Collibra offers flexible benefits designed to support employees and their loved ones through diverse circumstances and life events, including competitive compensation, health coverage, time off, and additional flexible offerings. Collibra is an equal opportunity employer and provides accommodations for applicants.

More jobs at Collibra

Similar jobs