
AI Infrastructure Software Engineer — CosmosLab
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.
Are you excited to explore new frontiers in AI? Join NVIDIA’s Cosmos Lab Infra team and take part in the innovation building the training infrastructure that supports our Physical AI world foundation models. Here is your opportunity to design, assemble, and improve the infrastructure for large-scale AI training, spanning pre-training, supervised fine-tuning (SFT), and reinforcement learning (RL) post-training. Embark on a journey where your work will be essential to influencing the future of AI!
What you'll be doing:
Create and implement the training infrastructure spanning pre-training, SFT, and RL post-training for Physical AI world foundation models. The work involves the framework and a comprehensive control plane across clusters to coordinate workloads efficiently.
Develop and improve the pre-training and SFT pipelines — large-scale data loading, distributed training, and checkpointing — to achieve high throughput and scalability.
Develop and improve the inference and evaluation stack, including the inference engine, inference/generation pipelines (which also support RL rollout), and evaluation pipelines. Use methods like continuous batching and KV-cache management to achieve high throughput and low latency.
Build and improve the effective interaction and data flow among the RL system's roles (policy, rollout, reward, simulation) while investigating system-level optimization opportunities.
Integrate and orchestrate simulation and robotics environments as RL environments — driving the simulation↔rollout↔training loop at scale.
Build and refine the distributed training backend — sharding/parallelism, mixed precision, activation checkpointing, and memory/throughput optimization across many GPUs.
Improve the efficiency, scalability, and resiliency of training and RL workloads — focusing on fault tolerance, fast/elastic restart, and throughput optimization under preemption and hardware failure.
Define meaningful, actionable reliability and efficiency metrics to track and improve system reliability.
Root cause, triage, and resolve failures from the application level down to the framework, GPU, and network/hardware level.
What we need to see:
5+ years developing software infrastructure for large-scale AI or distributed systems.
Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience).
Strong debugging and triage skills across the stack — from AI application down to GPU/hardware behavior.
Proven track record building and scaling large-scale distributed systems, ideally distributed training or inference.
Hands-on experience with AI training and/or inference infrastructure — RL/post-training, training frameworks, or inference serving.
Proficiency in Python (plus scripting), and solid software engineering practices: testing, defensive programming, version control, and CI.
Excellent communication and collaboration skills; intellectual curiosity, problem-solving, and willingness.
Ways to stand out from the crowd:
Experience building RL / post-training infrastructure — PPO/GRPO/DPO pipelines, rollout engines, and asynchronous RL.
Background with building large-scale, production-grade pre-training / SFT infrastructure.
Experience integrating simulation / robotics environments into training or RL loops — including vectorized environments and sim-to-real workflows.
Comprehensive knowledge of DL framework internals — PyTorch (FSDP/DTensor) and Megatron or equivalent experience, distributed training, and related optimization techniques.
Proficiency in C/C++/CUDA for performance-critical components and custom kernels.
Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/
Similar roles
Trainee Test Engineer
Side·Hyderabad, India
TRAINEE TEST ENGINEER – MANUAL GAME TESTING We are looking for a highly motivated Trainee Test Engineer to join our QA team in Hyderabad. This role provides an opportunity to work on manual game testing and be part of the quality assurance process, helping ensure the quality and functionality of games across various platforms. KEY RESPONSIBILITIES: * Conduct thorough testing…
- Contract
Senior Software Engineer, D2C
sonymusicentertainment·London, Canada
<p>At Sony Music Entertainment, we fuel the creative journey. We’ve played a pioneering role in music history, from the first-ever music label to the invention of the flat disc record. We’ve nurtured some of music’s most iconic artists and produced some of the most influential recordings of all time.</p> <p>Today, we work in more than 70 countries, supporting a diverse…
- On-site
- Full-time
- Express Entry — PR day one, no employer
Senior Software Engineer I (Agentic Engineering, Europe)
Coder·United Kingdom
As a Senior Software Engineer on Coder’s Agentic Engineering team, you’ll build and evolve the systems behind our agentic development experience. You’ll work across the agent harness, integrations, and workflows that connect agents with real development environments. You’ll stay hands-on, solve complex technical problems, and work closely with Product, Design, and other engineers to ship reliable agentic experiences. What you’ll…
- On-site
- Full-time
Dev/Ops Platform Engineer (m/w/d)
Klaus Scholz Personalvermittlung·Dormagen, Germany
AUFGABEN Ihre Aufgaben * Cloud Plattform gestalten: Sie bauen eine moderne Azure Cloud und Container Umgebung für geschäftskritische, selbst entwickelte Anwendungen auf und entwickeln diese kontinuierlich weiter * Migration vorantreiben: Sie begleiten die Transformation bestehender Docker und On Premises Umgebungen in eine skalierbare und zukunftsfähige Azure Infrastruktur * Container orchestrieren: Sie betreiben und optimieren containerisierte Anwendungen und arbeiten dabei mit…
- On-site
- Full-time
- Blue Card — tied; settle 21–33mo
Cloud Engineer (m/w/d) Azure
Klaus Scholz Personalvermittlung·Dormagen, Germany
AUFGABEN Ihre Aufgaben * Azure Plattform gestalten: Sie entwickeln die bestehende Microsoft Azure Infrastruktur weiter und sorgen dafür, dass Cloud Services zuverlässig, sicher und skalierbar zur Verfügung stehen * Cloud Infrastruktur aufbauen: Sie konzipieren und implementieren neue Cloud Komponenten und begleiten die Migration bestehender Systeme und Anwendungen in die Azure Umgebung * Container Plattform betreiben: Sie arbeiten mit Kubernetes, Azure…
- On-site
- Full-time
- Blue Card — tied; settle 21–33mo
AI Augmented Software Engineer [gn] Data Intelligence Platform
Jobgether·United Kingdom
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Augmented Software Engineer [gn] Data Intelligence Platform based in United Kingdom. This is an opportunity for a full-stack engineer to help build a modern data intelligence platform that enables organisations to discover, govern, and trust their…
- On-site
- Full-time
Engineering Manager (f/m/d)
Moonfare·Berlin, Germany
<div class="job__description body"> <p><strong>Join the team rewriting the rules in private markets.</strong></p> <p>Moonfare delivers what few others can: the highly sought-after funds and hidden-gem investments that go beyond what most private banks offer. Every opportunity is subjected to a ruthless vetting process; the bar is unforgivingly high. The result? Institutional-quality portfolios for investors who demand more.</p> <p>Our team combines finance…
- On-site
- Full-time
- Blue Card — tied; settle 21–33mo
Frontend Engineer
Omnea·London, Canada
OUR MISSION At Omnea, we’re reinventing how enterprise businesses operate, starting with the most painful parts: procurement – where a single purchase can drag on for months, trigger 50+ emails, and pull in Finance, Legal, Security, and IT just to get something approved. We’ve raised $75M from Khosla Ventures, Insight Partners, and Accel to change that. Our AI-native platform connects…
- On-site
- Full-time
- Express Entry — PR day one, no employer