
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
NVIDIA's Deep Learning Frameworks (DLFW) Infrastructure team is looking for a deeply technical Senior HPC Cluster Administrator to lead the design, deployment, and reliability of our large-scale GPU compute clusters. These systems run the most demanding deep learning training, inference, and high-performance computing workloads in the industry — from DGX/HGX platforms to ground-breaking Grace Blackwell systems. You will drive architectural decisions across compute, networking, and storage, and partner closely with software, research, and product teams to keep our infrastructure ahead of the workloads it supports.
What you'll be doing:
Own the full lifecycle of GPU compute clusters — procurement, provisioning, configuration management, monitoring, and deprecation — across heterogeneous Linux environments (DGX, HGX, embedded systems)
Design and scale storage solutions (NFS, Lustre, WekaFS, or equivalent) with a clear roadmap for capacity and performance growth
Lead automation of infrastructure using modern IaC tools (Ansible, Terraform) and CI/CD pipelines (GitLab)
Manage and optimize job scheduling via Slurm, including fair-share policies, reservation management, and MIG/GPU partitioning strategies
Maintain and improve observability stacks (Prometheus, Grafana, DCGM) and drive proactive resolution of hardware and software incidents
Collaborate with ML engineers and software teams to tune cluster configuration for large-scale distributed training workloads
Evaluate and introduce new technologies — networking fabrics (InfiniBand, NVLink, EFA/RDMA), storage tiers, container runtimes — to improve performance and reliability
Mentor junior engineers and contribute to team-wide engineering standards
What we need to see:
BS/MS in CS, EE, CE, or equivalent hands-on experience
5+ years of experience deploying and administering large-scale HPC or ML training clusters
Deep expertise in Linux systems administration at scale
Strong scripting and automation skills in Python and/or bash
Hands-on experience with Slurm (scheduling, accounting, cgroup configuration)
Proficiency with configuration management and IaC (Ansible required; Terraform a plus)
Experience with container technologies (Docker, Apptainer/Singularity, Kubernetes)
Solid understanding of high-speed networking (InfiniBand, RoCE, RDMA, EFA)
Experience with distributed/parallel filesystems and storage architecture
Ability to own problems end-to-end and communicate clearly with engineering and management stakeholders
Ways to stand out from the crowd:
Experience with NVIDIA GPU infrastructure tools (DCGM, nvidia-smi, MIG, NVSwitch diagnostics)
Familiarity with cluster management platforms (Colossus, Bright Cluster Manager, xCAT, or similar)
Experience supporting large-scale distributed deep learning workloads (PyTorch, JAX, Megatron)
Knowledge of BMC/IPMI/Redfish for out-of-band management and hardware lifecycle
Background in MLOps tooling or ML platform engineering
Join our team of world-class engineers and be part of the groundbreaking work we do at NVIDIA. We are committed to encouraging a collaborative and inclusive environment, where every team member has the opportunity to thrive and make a significant impact!
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. For Poland: The base salary range is 221,250 PLN - 383,500 PLN for Level 3, and 292,500 PLN - 507,000 PLN for Level 4.


Similar roles
Data Scientist, Growth
Lovable·London, Canada
TL;DR — The first data scientist in our new London growth team. You own the growth funnel end to end, from acquisition to activation to retention to expansion, find the levers, and work directly with growth engineering to ship experiments into the product. Not decision support: you drive the numbers. WHY LOVABLE? Lovable lets anyone and everyone build software with…
- On-site
- Full-time
- Express Entry — PR day one, no employer
Alts Data Ops Analyst
Addepar1·Edinburgh, United Kingdom
<div class="content-intro"><p><span style="text-decoration: underline;"><strong>Who We Are</strong></span></p> <p>Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. Addepar unifies portfolio, market and client data in a total portfolio view and delivers AI-powered insights within investment and client workflows. More than 1,400 firms in nearly 60 countries use Addepar to manage and advise…
- On-site
- Full-time
InnoMaster Big Data Engineering - part-time Master's Program
Innogames·Hamburg, Germany
Part-time Master's Program starting Oct 2026 or April 2027 Do you have a passion for Big Data? Do you want to make your Master studies exciting and practical? Do you wish to continue studying after your Bachelor's degree while simultaneously stepping into the professional world? And all this with a fixed monthly salary and complete coverage of tuition fees? Then,…
- On-site
- Full-time
- Blue Card — tied; settle 21–33mo
Data Engineer (all genders)
Gropyus·Berlin, Germany
<div class="content-intro"><p><strong>About The Company</strong></p> <p>GROPYUS is a technology-based construction company focused on building multi-story residential buildings. Thanks to its prefabricated building system with various design options, industrial offsite construction, and fully digitalized processes, the company manufactures aspirational, sustainable, and affordable homes using timber construction methods. GROPYUS is using scalable construction and manufacturing solutions to tap into a future market, boost…
- On-site
- Full-time
- Blue Card — tied; settle 21–33mo
Data Analyst: Retail Media
Constructor·Worldwide
ABOUT US Constructor is the next-generation platform for search and discovery in ecommerce, built to explicitly optimize for metrics like revenue, conversion rate, and profit. Our search engine is entirely invented in-house utilizing transformers and generative LLMs, and we use its core and personalization capabilities to power everything from search itself to recommendations to shopping agents. Engineering is by far…
- On-site
- Full-time
Senior Backend Engineer: Machine Learning Infrastructure
Constructor·Worldwide
ABOUT US Launched in 2019, Constructor is an AI-first e-commerce search and discovery platform that helps shoppers find the right products at the right time and enables leading global e-commerce brands to drive meaningful revenue and conversion gains. ABOUT THE ROLE The ML Infrastructure team builds and operates the shared backend services and platform capabilities that Constructor's ML and product…
- On-site
- Full-time
Working Student - Market & Business Intelligence (f/m/x)
Ionity Gmbh·Worldwide
Your mission Our Market Intelligence and Business Development team plays a pivotal role in driving our company's strategic growth within the fast-evolving eMobility sector. We focus on harnessing the power of data to inform our decision-making processes and guide our expansion. As part of our team, you'll develop sophisticated data models, enhance our internal databases, and deliver actionable insights that…
- On-site
- Full-time
Przedstawiciel handlowy / Przedstawicielka handlowa
MEGATEX AUTO sp. z o.o.·Imielin, Poland
Przedstawiciel handlowy / Przedstawicielka handlowa Miejsce pracy: Imielin Twój zakres obowiązków aktywne pozyskiwanie nowych klientów w segmencie akumulatorów przemysłowych i trakcyjnych, analiza rynku oraz identyfikacja potrzeb klientów, prowadzenie procesu sprzedaży od pierwszego kontaktu aż do finalizacji zamówienia, przygotowywanie ofert handlowych i wycen, dobór odpowiednich rozwiązań oraz ustalanie specyfikacji technicznej produktów zgodnie z potrzebami klienta, prowadzenie negocjacji ha…
- Full-time