Tech Stack
Tag name is followed by "@" symbol and proficiency level value.
About proficiency levels:
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
AI
API @ 4
HPC @ 7
IaC
Linux @ 7
Perl @ 4
Python @ 4
Slurm @ 4
- 1-2 — basic awareness. Minimal hands-on experience, and a rudimentary understanding of the technology's purpose;
- 3-6 — daily use. Comfortable and regular usage, capable of handling common tasks and challenges related to the technology;
- 7-9 — you are an expert, you can teach others, you know all the pitfalls and tricks;
- 10 — exceptional knowledge, comprehensive understanding, and adeptness in all aspects of the technology, including advanced problem-solving. Think twice before claiming or demanding such level.
Details
NVIDIA's silicon does not tape out without the farm behind it. Its EDA compute environment runs millions of cores across federated LSF cells, and every simulation, synthesis run, and timing signoff on its roadmap passes through it. NVIDIA is consolidating a dual-scheduler estate onto a single LSF platform and is seeking an engineer with deep knowledge of LSF internals, beyond configuration files.
This is a deep-specialist role responsible for resolving scheduling latency and other complex issues that are not explained by standard logs.
Responsibilities
- Own scheduler behavior across 15–25 federated LSF cells, including
mbatchdandmbschdtuning, scheduling cycle analysis, and contention patterns that appear as cells approach host-count ceilings. - Diagnose MultiCluster forwarding problems, including remote queue sizing, forwarding policy, and cross-cluster pending behavior.
- Define the technical design for cell topology and federation as the farm grows, including decisions about what belongs in a cell versus a new cell.
- Work with the IaC engineer to encode scheduler policy into a configuration schema that works reliably with MultiCluster at scale.
- Partner with CAD and methodology teams on workloads involving 500GB+ memory jobs, interactive-versus-batch contention, and tape-out crunch bursts.
Requirements
- Bachelor's or master's degree in Computer Science, Computer Engineering, or equivalent experience.
- 8+ years of experience in HPC or large-scale batch computing, including 5+ years working with IBM Spectrum LSF.
- Demonstrated depth in LSF internals, including the ability to debug scheduler behavior beyond the documentation and explain a scheduling cycle from submission to dispatch.
- Hands-on MultiCluster experience in a production, multi-site environment.
- Strong Linux systems fundamentals.
- Experience with system programming languages and scripting in Python, Perl, and shell.
Preferred Qualifications
- Experience working on LSF as a developer or in escalation engineering rather than only as a consumer of LSF.
- Experience with LSF integration points such as
esub,eexec,elim, submit wrappers, RTM, or LSF APIs. - Background in semiconductor or EDA computing, including environments where license constraints and job constraints compete.
- Experience migrating a production estate from Slurm, PBS, or Grid Engine without a scheduled outage noticed by users.
Compensation and Benefits
The base salary is determined based on location, experience, and the pay of employees in similar positions. The base salary ranges are:
- Level 4: $184,000–$287,500 USD per year
- Level 5: $224,000–$356,500 USD per year
The role also includes eligibility for equity and benefits.
Applications will be accepted at least until September 13, 2026. This posting is for an existing vacancy. NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer.