Senior Staff Software Engineer - Enterprise AI Platform

at Nvidia
USD 200,000-322,000 per year
SENIOR
✅ On-site

Tech Stack

AI @ 4 API @ 4 CI/CD @ 4 Codex Distributed Systems @ 8 Go @ 6 Observability Python @ 6 Rust @ 6 Security @ 4

Details

We are building an enterprise platform that enables long-running autonomous agents to operate safely at scale on shared accelerated compute. The platform provides agent blueprints, tools, skills, plugins, runtime safety, credential management, sandboxing, memory, and observability for teams across NVIDIA.

The platform abstracts differences between agent harnesses such as Claude, Codex, and open-source harnesses, allowing skills, safety policies, and traces to work consistently across execution models. It supports governed multi-agent collaboration, delegation, handoffs, identity and policy propagation, live evaluation signals, and continuous improvement based on production telemetry.

Responsibilities

  • Design agent blueprints with interfaces for authorization, sandboxing, memory, observability, and skills.
  • Build a runtime safety harness with policy enforcement, rate and budget caps, circuit breakers, approval gates, action allow-lists, and kill switches.
  • Enable agent composition and orchestration through declarative skill manifests, delegation, handoffs, and support for headless, long-running autonomous agents.
  • Broker credentials for multi-agent systems using least-privilege token scoping without exposing secrets.
  • Provide checkpointing and recovery so agents can resume after crashes or restarts.
  • Build cross-harness observability and evaluation, including decision-level traces, correlation IDs, cost anomaly alerts, and quality scoring across skills, agents, and products.
  • Turn evaluation and production signals into insights that improve agent behavior.
  • Participate in an on-call rotation to maintain the health and reliability of the production platform.

Requirements

  • BS or MS in Computer Science, Engineering, or a related field, or equivalent experience.
  • 12+ years of experience building distributed systems, infrastructure, or developer platforms at scale.
  • Hands-on experience building agents on an agent harness, exposing them as APIs, and deploying them with CI/CD.
  • Experience deploying and operating long-running services on container orchestration platforms.
  • Experience with scalable-system building blocks such as messaging, caching, and durable storage.
  • Proficiency in Python, Go, Rust, or a similar programming language.

Preferred Qualifications

  • Experience building safety or policy engines that enforce runtime rules on agent actions, including approval gates and kill switches.
  • Experience designing evaluation and feedback loops for agent behavior tied to versioned skills or blueprints.
  • Experience building self-evolving loops that improve agents using evaluation and production signals.
  • Knowledge of security fundamentals including threat modeling, authentication, authorization, least privilege, secrets management, and token exchange.
  • Experience designing AI data platform components such as ingestion pipelines, vector stores, and retrieval APIs.
  • Experience shipping platform building blocks adopted by multiple engineering teams.
  • Experience leading complex technical projects, including migrations or greenfield platform builds, aligning teams, and writing clear design documentation.

Compensation and Benefits

The base salary range is USD 200,000–322,000, determined by location, experience, and the compensation of employees in similar positions. The role is also eligible for equity and benefits.

Applications will be accepted at least until September 13, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.

More jobs at Nvidia

Similar jobs