Senior Systems Software Engineer, Developer Productivity And Cloud Automation - GeForce NOW
Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
API @ 4
AWS
Ansible @ 4
Azure
CI/CD @ 4
Compliance
Datadog
DevOps @ 8
GCP
Go @ 6
Grafana
Helm @ 4
Jenkins @ 6
Kubernetes @ 8
Microservices
Observability @ 4
Prometheus
Python @ 6
Security @ 6
Slack
Terraform @ 4
Vault @ 4
gRPC @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Role description
GeForce NOW is NVIDIA's Cloud Gaming service, streaming games at the highest quality to any and every user, regardless of their device type and capabilities. GeForce NOW uses NVIDIA proprietary software and always-the-latest hardware to enable a near-instant launch experience.
Our mission empowers every GeForce NOW engineer to deliver confidently by removing obstacles, catching issues early, and turning deployments into a strength rather than a liability. GeForce NOW handles pipelines using GitLab's CI platform and Jenkins, along with Flux CD and Argo CD rollouts, StackStorm event-driven automation, HashiCorp Vault for credential storage, and automation across multi-cloud environments.
This role is for a Kubernetes-native deployment platform paired with its automation system, and the backend services and APIs that support it. The team delivers zero-downtime global rollouts with automatic drift detection, provides reliable staging environments replicating production accurately, and builds cloud automation for bare-metal and multi-cloud environments. If a GFN service ships, runs, or self-heals, you build the infrastructure behind it.
Responsibilities
- Build and develop backend microservices and REST/gRPC/MCP APIs powering the deployment platform, zone reservation/lease system, and developer self-service tooling
- Extend the platform with dynamic delivery, automatic rollback, drift detection, and automated zone bootstrapping
- Build and stabilize prod-like staging environments/zones so teams catch regressions early
- Develop and sustain GitOps pipelines (Flux CD, Argo CD) across on-premises and Nvidia GFN Cloud/AWS/Azure/GCP
- Develop Kubernetes CRDs and operators in Go for scheduling, auto-scaling, and compliance across data centers
- Build backend integrations and control-plane services connecting CI/CD, observability, and automation systems into a unified platform experience
- Automate dedicated hardware and multiple cloud platform configurations using Terraform, Ansible, and Vault
- Implement monitoring solutions including Prometheus, Grafana, Datadog, and ELK, paired with SLO/alerting for early detection
- Integrate automation tools — runbooks, StackStorm bots, anomaly-triggered remediation, and Slack self-service release bots
Requirements
- Bachelor or higher degree in computer science, engineering, or equivalent experience
- 10+ years of experience in Cloud Infrastructure and DevOps, with deep expertise in Kubernetes, GitOps (or equivalent), and production-grade cloud-native CI/CD pipelines
- Expert Kubernetes: CRDs, operators, multi-cluster management, and security hardening (CIS, PCI/SOC 2)
- Proficiency in Flux CD or Argo CD, GitLab CI, and Jenkins; Go or Python for control plane development
- Experience handling Vault, Terraform, Ansible, and Helm in both on-premises and cloud environments
- Experience developing and scaling RESTful, gRPC, MCP APIs and backend services
- Experience running hybrid multi-cloud and bare-metal at production scale, and owning a platform roadmap end-to-end
Benefits
- Competitive salaries and a generous benefits package
- You will also be eligible for equity and benefits
Applications for this job will be accepted at least until July 30, 2026.