
DL Performance Software Engineer - LLM Inference
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference software, optimize GPU kernels, drive industry benchmarks, and work with state-of-the-art research techniques to improve serving efficiency. You’ll collaborate across inference performance, kernels, training, large-scale serving, and research teams to push the frontier of accelerated computing for AI.
What you’ll be doing:
Contribute features to vLLM that empower the newest models with the latest NVIDIA GPU hardware features and serving runtime algorithms.
Profile and optimize the inference framework (vLLM) with methods like speculative decoding, 5D Parallelism, and prefill-decode disaggregation.
Architect novel frameworks and runtime optimizations for inference infrastructure, benchmarking, and kernels.
Conduct and publish original research that advances the Pareto frontier in ML Systems; survey recent publications and find a way to integrate research ideas and prototypes into production-grade, open-source software.
Develop, optimize, and benchmark GPU kernels (both hand-tuned and compiler-generated) using techniques such as fusion, autotuning, and memory/layout optimization.
What we need to see:
Bachelor’s, Master’s, or PhD degree in Computer Science (CS), Computer Engineering (CE) or Software Engineering (SE).
5+ years of industry experience in software engineering or equivalent research experience.
Strong programming skills in Python and one of C/C++, Go, or Rust. Solid CS fundamentals: algorithms & data structures, operating systems, computer architecture, parallel programming, software engineering, distributed systems, deep learning theories.
Knowledgeable and passionate about performance engineering in ML frameworks (e.g., PyTorch) and inference engines (e.g., vLLM and SGLang).
Familiarity with GPU programming and performance: CUDA, memory hierarchy, streams, NCCL; proficiency with profiling/debug tools (e.g., Nsight Systems/Compute).
Excellent debugging, problem-solving, and communication skills; ability to excel in a fast-paced, multi-functional setting.
Ways to Stand out from the Crowd
Experience developing major features and optimizations for LLM inference engines (e.g., vLLM, SGLang).
Hands-on work with LLM inference and training runtimes (deploying LLMs to production, large-scale LLM pre-training and RL), ML compilers and DSLs (e.g., Triton, CuTe, MLIR/LLVM, XLA), GPU libraries (e.g., CUTLASS) and features (e.g., CUDA Graph, Tensor Cores).
Experience with speculative decoding training and runtime features: tree-structured drafting, parallel drafting, diffusion LLMs, DFlash, EAGLE.
Contributions to open-source projects and/or publications; please include links to GitHub pull requests, published papers and artifacts.
At NVIDIA, we believe artificial intelligence (AI) will fundamentally transform how people live and work. Our mission is to advance AI research and development to create groundbreaking technologies that enable anyone to harness the power of AI and benefit from its potential.
Our team consists of experts in AI, systems and performance optimization. Our leadership includes world-renowned experts in AI systems who have received multiple academic and industry research awards. If you’re excited to build systems, kernels, and tools that make large-scale AI faster, more efficient, and easier to deploy, we’d love to hear from you.
#LI-Hybrid
You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until August 10, 2026.This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
Similar roles
Staff Software Engineer, Time and Scheduling
Gusto, Inc.·Worldwide
---------------------------------------- About Gusto At Gusto, we're on a mission to grow the small business economy. We handle the hard stuff — payroll, health insurance, 401(k)s, and HR — so owners can focus on their craft and their customers. With teams in Denver, San Francisco, and New York, we support more than 500,000 small businesses nationwide and are building a workplace…
- Remote
- Full-time
Senior Site Reliability Engineer (Performance and Scalability)
Digital Zone·United Arab Emirates
Your mission is to make DigitalZone able to scale. You will build the platform's capacity to absorb campaign-level traffic spikes, and you will give every engineering team the tools, standards, and practices to load- and failure test their own systems. This is an enablement role at its core: you raise the reliability bar across the org by building capability, not…
- Remote
- Full-time
- Green Visa — self-sponsored, no tie
Software Developer - Risk Technology
Squarepointcapital·London, United Kingdom
<p><strong>Position Overview:</strong></p> <p>Risk Technology is a global team that designs, builds and maintains Squarepoint’s trading risk platform, which is responsible for trade capture, position management, profit/loss computation, inventory/locate management and internal order routing. These critical systems need to be performant, resilient, and capable of timely processing of high volumes of trading data in both live and historical scenarios, requiring solutions…
- On-site
- Full-time
Software Developer - Data Pipelines (Python)
Squarepointcapital·London, United Kingdom
<p><strong>Position Overview:</strong></p> <p>We are seeking an experienced Python developer to join our Alpha Data team, responsible for delivering a vast quantity of data served to users worldwide. You will be a cornerstone of a growing Data team, becoming a technical subject matter expert and developing strong working relationships with quant researchers, traders, and fellow colleagues across our Technology organization.</p> <p>Alpha…
- On-site
- Full-time
Junior Software Developer - Front-end
Squarepointcapital·London, United Kingdom
<p><strong>Please only apply to the one job you feel best fits your skillset and experience. If our team feels you are better suited for another role, we will reach out about the alternate opportunity.</strong></p> <p><strong>Position Overview:</strong></p> <p><span class="ui-provider a b c d e f g h i j k l m n o p q r s t u v…
- On-site
- Full-time
Junior QA Engineer
AoFrio·New Zealand
Description Welcome to our World of Cold! At AoFrio, we are global leaders in providing IoT solutions to the food and beverage industry. Our innovative technology and dedicated team have positioned us at the forefront of our market. We are proud leaders in hardware-enabled Software as a Service (SaaS) for commercial refrigeration, with cutting-edge IoT solutions used by major brands…
- Full-time
- Skilled Migrant (Residence) — PR, points-based, no single employer
Software Engineer Specialist – Integration & AI
AIA Group·New Zealand
Your Role with Us As a Software Engineer Specialist – Integration in our Technology team, you’ll play a key role in designing and delivering integration solutions that keep our systems connected, secure, and running smoothly. You’ll collaborate with architects, engineers, and business stakeholders to build robust, scalable platforms that support both strategic initiatives and day-to-day operations. This is an exciting…
- Full-time
- Skilled Migrant (Residence) — PR, points-based, no single employer
(Intern) AI Automation Engineer
Tacto·Munich, Germany
YOUR IMPACT As part of Solution Engineering at Tacto, you'll bridge our powerful platform with measurable customer value through data expertise. Working directly with customers and leads, you'll understand their unique supply chain challenges and implement tailored technical solutions. By integrating customer procurement data and configuring the platform to match their processes, you'll drive adoption and showcase immediate ROI (even…
- On-site
- Internship
- Blue Card — tied; settle 21–33mo