
Engineering Manager, Kubernetes Customer Delivery and Self-Service
NVIDIA’s DGX Cloud Kubernetes Platform & Production Engineering team is seeking an Engineering Manager to develop and guide our Customer Delivery and Self-Service function. This group will manage the entire engineering delivery process from an approved customer request through platform enablement, cluster build, validation, and handoff. The leader will also transform the current cross-team process into a scalable, automated, self-service solution.
What you’ll be doing:
Build and lead a team of software and production engineers passionate about Kubernetes customer delivery, onboarding, and self-service.
Own end-to-end delivery of production Kubernetes clusters for AI workloads, from accepted request through enablement, qualification, validation, and customer handoff.
Drive a coordinated delivery plan with clear owners, dependencies, readiness gates, timelines, risks, status, and blocking issues.
Partner across platform, runtime, release, fleet operations, CSE, product, TPM, security, and infrastructure teams.
Build integrations connecting customer intake and status systems with Kubernetes provisioning, access, validation, and production acceptance.
Turn recurring delivery tasks into detailed, automated self-service workflows using APIs, AI tools, and agents.
Define service interfaces and measure and improve delivery speed, readiness, automation, recovery, and customer visibility.
Set the team’s roadmap, staffing, and operational ownership while hiring, mentoring, and developing technical leaders.
What we need to see:
8+ overall years of industry experience, including 2+ years leading or managing engineers.
Experience building platform APIs, self-service infrastructure, workflow automation, developer platforms, or customer onboarding systems.
Strong understanding of Kubernetes, cloud infrastructure, distributed systems, or production engineering.
Hands-on experience using AI coding tools and AI-enabled engineering workflows.
Experience integrating multiple systems and teams into a reliable end-to-end workflow.
Ability to translate customer and operational requirements into clear technical interfaces and automated solutions.
Strong cross-functional leadership, communication, customer empathy, prioritization, and judgment.
BS or MS in Computer Science, Engineering, or equivalent experience.
Ways to stand out from the crowd:
Experience building Kubernetes provisioning, infrastructure-as-code, service catalog, or internal developer platform capabilities.
Familiarity with Terraform, GitOps, identity and access management, RBAC, APIs, workflow engines, and production-readiness automation.
Experience with GPU infrastructure and AI-optimized Kubernetes clusters, including accelerated networking, high-performance storage, GPU scheduling, workload qualification, or large-scale fleet operations.
A track record of reducing onboarding time and operational toil through automation and self-service.
Experience combining strong platform engineering with an attitude centered on product development and customer needs.
Join us in transforming the future of computing and make an impact on the world!
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD.You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until August 15, 2026.This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.Similar roles
Software Engineer, Atlas Distributed Systems
Rubrik Job Board·Palo Alto, United States
ABOUT TEAM Rubrik Atlas is the core data path for all Rubrik products, whether in the data center, at the edge, or in the cloud. It is a distributed, scale-out, fault tolerant, performant, deduplicated user-space filesystem that uniquely combines block device (ext4/NFS/SMB) and S3 backend storage interfaces. It has been the cornerstone of Rubrik’s innovative tech stack since Day 1.…
- Full-time
QA Engineer
veritree·Vancouver, Canada
VERITREE AND JOB OVERVIEW veritree is an award-winning climate tech start-up based in Vancouver. Launched in 2021, our technology measures and verifies the impact of global restoration efforts from the ground up. We are on a mission to plant 1 billion verified trees by 2030, collaborating with businesses, planting organizations, and consumers who believe in the transformative power of verified…
- Hybrid
- Full-time
- Express Entry — PR day one, no employer
WerkstudentIn AI Agent Developer
logen.ai·Berlin, Germany
Du hast Erfahrung mit KI und willst sie an echten Kundenprojekten einsetzen? Bei uns baust du AI Agents, die in Produktion gehen. Keine Proof-of-Concepts für die Schublade. Wir sind ein KI-Startup aus Berlin, spezialisiert auf Automatisierung im Kundenservice. Wir bauen AI Agents, Voicebots und Automatisierungslösungen für Unternehmen im deutschen Mittelstand. Gegründet 2023 von Oscar Schwarz und Ludwig Sickert, aktuell ein…
- On-site
- Full-time
- Blue Card — tied; settle 21–33mo
Test Engineer (Java/.Net)
Solirius Reply·London, United Kingdom
About Us: Solirius Reply, part of the Reply Group, is a technology consultancy and digital transformation partner that helps organisations solve complex challenges through strategy, design, engineering, and delivery. We work closely with our clients to deliver secure, accessible, user-focused services that evolve with their needs. By combining deep technical expertise with people-centred design, we create solutions that deliver meaningful,…
- Hybrid
- Full-time
Platform Engineer (Mid-level)
Team17 Digital·Wakefield, United Kingdom
About the Role We are seeking a Platform Engineer (Mid-level) to join our Platform Engineering function. This hands-on technical role focuses on the platforms, infrastructure, automation, and tooling that support our development teams and business operations. The successful candidate will help maintain, improve, and modernise our platform estate across cloud services, infrastructure, monitoring, automation, CI/CD, and operational tooling. The role…
- Full-time
Deployment Lead - ERP Solutions
Rillet·Worldwide
WHAT WE DO Rillet serves accounting and finance teams, the financial brains of their companies. Our job is to help them run the numbers with impossible speed, accuracy, and insight. Rillet is the AI-native ERP built for modern finance teams. We help companies automate the most critical parts of finance, from closing the books and recognizing revenue to consolidating entities…
- On-site
- Full-time
Solutions Consultant
Rillet·Worldwide
WHAT WE DO Rillet serves accounting and finance teams, the financial brains of their companies. Our job is to help them run the numbers with impossible speed, accuracy, and insight. Rillet is the AI-native ERP built for modern finance teams. We help companies automate the most critical parts of finance, from closing the books and recognizing revenue to consolidating entities…
- On-site
- Full-time
Accounting Solutions Consultant
Rillet·Worldwide
WHAT WE DO Rillet serves accounting and finance teams, the financial brains of their companies. Our job is to help them run the numbers with impossible speed, accuracy, and insight. Rillet is the AI-native ERP built for modern finance teams. We help companies automate the most critical parts of finance, from closing the books and recognizing revenue to consolidating entities…
- On-site
- Full-time