Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 6
Communication @ 7
DevOps @ 7
GPU
GenAI
Generative AI @ 6
Kubernetes @ 4
Leadership @ 4
Linux @ 7
Mentoring @ 4
Networking @ 7
Observability @ 3
SRE
Security
Software Development @ 7
Technical Leadership @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
As a Senior DevOps Engineer, you will help lead the evolution of infrastructure operations within the Networking Software group. You will build, automate, and operate scalable platforms that support networking software development and testing, using a strong Linux systems administration foundation. The role involves technical leadership, reliability engineering, automation, and continuous improvement across global engineering sites.
Responsibilities
- Build, provision, configure, and maintain scalable Linux infrastructure for networking feature creation and validation, including physical servers, network switches, virtualization platforms, containers, and remote-management interfaces.
- Develop automation for infrastructure provisioning, configuration management, software deployment, upgrades, and day-to-day operations using infrastructure-as-code and configuration-management practices.
- Build reusable tools and self-service capabilities that simplify infrastructure operations, improve engineering efficiency, and reduce repetitive manual work.
- Diagnose and resolve complex issues spanning hardware, firmware, operating systems, virtualization, containers, storage, network communications, and application environments.
- Implement monitoring, observability, capacity management, and reliability practices to improve infrastructure performance, availability, and operational readiness.
- Partner with engineering, IT, facilities, security, and network teams to define technical standards, maintain documentation and runbooks, and establish scalable operational processes.
- Provide technical leadership, guide infrastructure initiatives, and mentor team members in automation, troubleshooting, and operational guidelines.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or equivalent experience.
- At least 6 years of experience in systems engineering, DevOps, site reliability engineering, or infrastructure operations, including significant hands-on experience coordinating production or engineering Linux environments.
- Experience working in the semiconductor industry or a hardware-focused engineering environment, with hands-on expertise in bare-metal Linux systems and server components, including CPUs, GPUs, memory, PCIe devices, NICs, storage, BIOS/UEFI, BMC/IPMI/Redfish, power, and cooling.
- Proficiency in generative AI tools and the ability to apply them to automation, troubleshooting, documentation, operational analysis, and engineering efficiency.
- Ability to diagnose complex issues across hardware, firmware, and operating-system layers.
- Strong data-center networking knowledge, including TCP/IP, DNS, DHCP, VLANs, routing, switching, firewalls, and network troubleshooting tools.
- Strong analytical, problem-solving, written communication, and cross-departmental collaboration skills, with the ability to guide technical initiatives to completion.
Preferred Qualifications
- Experience managing Linux KVM/QEMU virtualization, Kubernetes clusters, multi-user engineering lab environments, NFS or distributed storage systems, and automated OS or cluster provisioning platforms.
- Experience supporting fast-growing engineering labs, large-scale data-center environments, or globally distributed infrastructure.
- Familiarity with observability platforms, including metrics, logging, tracing, alerting, incident management, and service-level objectives.
- Experience creating self-service infrastructure platforms and reusable automation that improves developer efficiency.
- Ability to establish clear, reliable, and scalable engineering and operational practices from evolving requirements.
- Proven technical leadership, team leadership, or managerial experience, including mentoring engineers, prioritizing work, coordinating cross-functional initiatives, and driving projects from planning through completion.
Compensation and Benefits
The base salary range is $184,000–$287,500 for Level 4 and $224,000–$356,500 for Level 5. Salary will be determined based on location, experience, and the pay of employees in similar positions. The role also includes eligibility for equity and benefits.
Applications will be accepted at least until September 21, 2026. NVIDIA uses AI tools in its recruiting processes and is an equal opportunity employer.