Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI @ 4
API @ 6
Ansible @ 7
CI/CD @ 7
Communication @ 6
Compliance @ 4
Databricks @ 7
Debugging @ 6
Distributed Systems @ 8
Go @ 6
IaC
Java @ 6
Kubernetes @ 7
Leadership @ 6
Linux @ 7
Networking @ 4
Observability
Python @ 6
Salt @ 7
Security
Technical Leadership @ 6
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA is looking for a Principal Software Engineer to join its Configuration Management team and define the future of enterprise infrastructure automation, configuration management, and operational platforms. This deeply technical, hands-on leadership role will architect and build foundational software systems that manage infrastructure consistently across compute, storage, networking, data centers, cloud, and hybrid environments.
The role involves establishing multi-year technical direction, identifying organization-wide challenges, and driving large-scale transformations from strategy through architecture, implementation, production adoption, and measurable outcomes.
Responsibilities
- Define the multi-year technical vision and architecture for IT infrastructure automation, configuration management, orchestration, and self-service platforms.
- Set the technical direction for infrastructure automation and configuration management using technologies such as Ansible Automation Platform, AWX, Salt, or equivalent platforms.
- Establish strategies for applying AI to configuration management, including configuration intelligence, drift and compliance analysis, change-risk identification, root-cause assistance, intelligent recommendations, and guarded automated remediation.
- Develop prototypes and production software, review critical code and designs, and resolve challenging technical and scalability problems.
- Build secure and scalable integrations across infrastructure platforms, cloud services, CMDB, secrets management, observability, and enterprise data systems.
- Set and drive configuration management strategy across networking, storage, and compute domains.
- Establish architectures and engineering standards, and lead implementation and adoption within large-scale environments.
- Ensure platforms meet enterprise requirements for availability, scalability, performance, security, disaster recovery, observability, and operational support.
- Mentor senior and staff engineers, raise engineering standards, facilitate architectural decisions, and develop technical leaders across teams.
- Partner with engineering and executive leadership to translate critical business challenges into technical strategy, prioritized roadmaps, and measurable business outcomes.
Requirements
- 15+ years of progressive software engineering experience, with a sustained record of delivering complex, business-critical platforms and distributed systems.
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a similar domain, or equivalent experience.
- Deep knowledge of enterprise configuration management or infrastructure automation using Ansible Automation Platform, AWX, Salt, or equivalent platforms.
- Deep hands-on expertise in Go, Python, Java, or a comparable systems programming language, including APIs, concurrency, distributed systems, testing, debugging, and performance engineering.
- Strong hands-on experience with Linux, Kubernetes, containers, cloud and hybrid infrastructure, CI/CD, and Infrastructure as Code.
- Strong experience developing and implementing enterprise data and automation pipelines using Databricks or comparable large-scale data platforms.
- Experience defining, owning, and evolving the architecture of large-scale infrastructure platforms operating across multiple teams, data centers, or cloud environments.
- Ability to identify organization-wide business and technical challenges, establish a clear strategy, and drive implementation across teams.
- Outstanding communication and technical leadership skills, including the ability to build consensus, influence senior leaders, mentor experienced engineers, and lead through ambiguity.
Preferred Qualifications
- Experience architecting, deploying, and managing Ansible Automation Platform at enterprise scale.
- Experience developing platforms that manage large global infrastructure fleets across data centers, public clouds, compute, storage, and networking environments.
- Experience leading the development and enterprise-wide adoption of an AI-enabled configuration management, infrastructure automation, or autonomous operations platform.
- Experience using AI to solve infrastructure challenges such as configuration drift, compliance, change-risk analysis, incident diagnosis, capacity management, predictive operations, or automated remediation.
Benefits
NVIDIA offers competitive salaries, equity, and a comprehensive benefits package. The company is committed to fostering an inclusive work environment and is an equal opportunity employer.