Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
AWS @ 4
Ansible @ 4
ArgoCD @ 4
Azure @ 4
Bash @ 7
CI/CD @ 4
Dashboarding
DevOps @ 6
GCP @ 4
GPU @ 4
Grafana @ 4
Kubernetes @ 7
Linux @ 7
Networking @ 7
Observability @ 4
OpenTelemetry @ 4
Performance Analysis
Prometheus @ 4
Python @ 7
SRE @ 6
Security
Software Development @ 7
Terraform @ 4
gRPC @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
We are looking for a Senior System Software Engineer, Software Defined Networking, to design, build, and operate highly performant and scalable SDN solutions for NVIDIA's AI Clouds hosting GPU-accelerated workloads, including hyperscale multi-node training, inference, cloud gaming, and cloud functions.
This role spans the full lifecycle of the SDN stack, from designing and developing new control and data plane software to ensuring operational excellence in production through reliability engineering, CI/CD, observability, and incident response.
Responsibilities
- Design and develop next-generation multi-tenant cloud SDN control and data plane software using OVS, OVN, and OpenFlow.
- Build Infrastructure-as-a-Service virtual network orchestration and services using gRPC and REST to support tenant workload security and performance SLAs for BMaaS, VMaaS, and Kubernetes.
- Drive upstream contributions to OVN-Kubernetes and related open-source projects.
- Develop software for network observability, including monitoring, telemetry, intelligent metering, and performance analysis.
- Operate and support OVS-OVN-based SDN solutions in large-scale NVIDIA AI Cloud environments.
- Own end-to-end observability for the SDN stack by building and maintaining monitoring, alerting, distributed tracing, and dashboarding for real-time insight into network health, performance, and tenant SLAs.
- Design, enhance, and maintain GitLab CI/CD pipelines across Linux host networking, OVS, OVN, and Kubernetes CNIs.
- Implement GitOps approaches or related solutions for secure, seamless integration with cloud infrastructure.
- Drive reliability through incident management, resource monitoring, and performance tuning.
- Collaborate with SRE, DevOps, and network engineering teams on production readiness and operational tooling.
Requirements
- BS or MS in Computer Science or a related technical field, or equivalent experience.
- 8+ years of proven experience in software development for large-scale distributed environments.
- Expert-level knowledge of OVN, OVS, OpenFlow, and modern network protocols.
- Strong programming skills in C and Go, with advanced scripting skills in Bash and Python.
- Deep knowledge of Kubernetes and practical experience deploying and supporting CNIs, particularly OVN-Kubernetes.
- Hands-on experience with Infrastructure-as-Code and deployment tools such as Ansible, Terraform, ArgoCD, and Flux.
- Experience designing and operating complex, multi-stage CI/CD pipelines.
- Hands-on experience developing secure, high-performance services using gRPC and REST with TLS and strong authentication.
- Strong knowledge of datacenter routing, switching, and Linux host and VM networking.
Preferred Qualifications
- Contributions to open-source projects, especially OVS, OVN, OVN-Kubernetes, or other Kubernetes networking projects.
- Experience with hardware acceleration, including GPU, DPU, or equivalent networking technologies.
- Practical experience with major cloud providers such as AWS, Azure, and GCP, as well as hybrid and multi-cloud deployments.
- Advanced SRE or DevOps expertise, including on-call responsibilities, incident management, service reliability targets, and production ownership.
- Experience with observability platforms and tools such as Prometheus, Grafana, Jaeger, OpenTelemetry, and ELK.
Compensation and Benefits
The base salary range is USD 184,000–287,500. Base salary will be determined based on location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.
Applications will be accepted at least until August 18, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes and is committed to an inclusive, equal-opportunity work environment.