Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 4
Agentic AI @ 4
Automated Testing @ 4
Bash @ 4
CI/CD @ 4
Debugging @ 4
Distributed Systems @ 4
GPU
Kubernetes @ 4
Leadership @ 6
Linux @ 7
Mentoring @ 6
Networking @ 4
Observability @ 4
Performance Analysis @ 4
Profiling
Python @ 4
Security @ 7
Technical Leadership @ 6
gRPC @ 7
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
Within NVIDIA Gaming Network infrastructure and storage, the NGN Software-Defined Storage team develops storage and infrastructure software for NGN and GeForce NOW, including Kubernetes, block storage, performance and scale, security, tenant storage, and related central services.
NVIDIA is seeking a Senior System Software Engineer to advance software-defined networking for global GPU cloud infrastructure supporting AI workloads, cloud gaming, content delivery, and other accelerated services. The role involves building and operating secure, programmable cloud networking using Open vSwitch, OVN, Kubernetes, C, and Go. The engineer will collaborate with senior engineers, NVIDIA architects, and partner teams on architecture, implementation, qualification, issue response, and production improvements.
Responsibilities
- Collaborate with engineers and architects to define, review, and evolve multi-tenant control-plane and data-plane architecture using Open vSwitch, OVN, OpenFlow, and overlay network technologies.
- Provide technical leadership and mentorship through design and code reviews, implementation guidance, production-readiness decisions, and engineering-quality improvements.
- Develop and review production software for Kubernetes networking, network plugins, distributed control planes, Linux host networking, virtual-machine networking, and network automation.
- Translate product and infrastructure requirements into secure orchestration services, APIs, component boundaries, state models, compatibility strategies, and delivery plans.
- Drive Open vSwitch and OVN integration across flow behavior, configuration, lifecycle management, upgrades, interoperability, performance, failure recovery, and large-scale production operations.
- Build automated unit, integration, system, performance, scale, and upgrade tests connected to continuous integration, deployment, and release-qualification workflows.
- Participate in the on-call rotation and lead networking escalations with partner teams. Use packet captures, SDN state, telemetry, profiling, controlled experiments, and source debugging to resolve incidents and prevent recurrence.
- Define and improve reliability, performance, security, and resource-efficiency objectives, including monitoring, telemetry, tracing, and service-level visibility for production networks.
Requirements
- Bachelor's or master's degree in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- 12 or more years of experience designing, implementing, testing, and maintaining production software in both C and Go.
- Experience using Bash and Python for testing, diagnostics, builds, or operational automation.
- Hands-on experience developing, integrating, and troubleshooting Open vSwitch and OVN, including OpenFlow and control-plane-to-data-plane behavior.
- Production Kubernetes networking experience, including container network interfaces, network policy, node and pod traffic paths, network plugins, upgrades, and failure modes.
- Strong Linux networking fundamentals and practical knowledge of IP, TCP, UDP, routing, switching, overlay networks, tunneling, network namespaces, and network policy.
- Experience architecting distributed systems with secure service APIs, state management, consistency, scalability, compatibility, and failure handling.
- Experience with the production lifecycle, including automated testing, CI/CD, deployments, upgrades, observability, and performance analysis.
- Experience supporting production services through an on-call rotation and coordinating incident response using flows, logs, metrics, traces, profiles, experiments, and source-level debugging.
- Demonstrated collaborative architecture and technical leadership through written designs, significant technical decisions, cross-team work, critical reviews, mentoring, and delivery of systems into production.
Preferred Qualifications
- Contributions to Open vSwitch, OVN, OVN-Kubernetes, Kubernetes networking, or another relevant open-source project.
- Experience with large-scale cloud and accelerated virtualization systems, including SR-IOV, RDMA, SmartNICs, data processing units, network function virtualization, KVM/QEMU, or container runtimes.
- Experience building secure, high-performance gRPC or REST services using transport security and strong authentication.
- Experience using agentic AI and AI-assisted software-development tools, including coding agents, reusable skills, or Model Context Protocol integrations.
Benefits
The position offers equity and benefits in addition to the base salary. NVIDIA promotes an inclusive work environment and is an equal opportunity employer.