Nvidia brand banner

DL Performance Software Engineer - LLM Inference

NvidiaToronto, CanadaPosted 3d ago
via Workday

We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You’ll architect and implement high-performance inference software, optimize GPU kernels, drive industry benchmarks, and work with state-of-the-art research techniques to improve serving efficiency. You’ll collaborate across inference performance, kernels, training, large-scale serving, and research teams to push the frontier of accelerated computing for AI.
 

What you’ll be doing:

  • Contribute features to vLLM that empower the newest models with the latest NVIDIA GPU hardware features and serving runtime algorithms.

  • Profile and optimize the inference framework (vLLM) with methods like speculative decoding, 5D Parallelism, and prefill-decode disaggregation.

  • Architect novel frameworks and runtime optimizations for inference infrastructure, benchmarking, and kernels.

  • Conduct and publish original research that advances the Pareto frontier in ML Systems; survey recent publications and find a way to integrate research ideas and prototypes into production-grade, open-source software.

  • Develop, optimize, and benchmark GPU kernels (both hand-tuned and compiler-generated) using techniques such as fusion, autotuning, and memory/layout optimization.

What we need to see:

  • Bachelor’s, Master’s, or PhD degree in Computer Science (CS), Computer Engineering (CE) or Software Engineering (SE).

  • 5+ years of industry experience in software engineering or equivalent research experience. 

  • Strong programming skills in Python and one of C/C++, Go, or Rust. Solid CS fundamentals: algorithms & data structures, operating systems, computer architecture, parallel programming, software engineering, distributed systems, deep learning theories.

  • Knowledgeable and passionate about performance engineering in ML frameworks (e.g., PyTorch) and inference engines (e.g., vLLM and SGLang).

  • Familiarity with GPU programming and performance: CUDA, memory hierarchy, streams, NCCL; proficiency with profiling/debug tools (e.g., Nsight Systems/Compute).

  • Excellent debugging, problem-solving, and communication skills; ability to excel in a fast-paced, multi-functional setting.

Ways to Stand out from the Crowd

  • Experience developing major features and optimizations for LLM inference engines (e.g., vLLM, SGLang).

  • Hands-on work with LLM inference and training runtimes (deploying LLMs to production, large-scale LLM pre-training and RL), ML compilers and DSLs (e.g., Triton, CuTe, MLIR/LLVM, XLA), GPU libraries (e.g., CUTLASS) and features (e.g., CUDA Graph, Tensor Cores).

  • Experience with speculative decoding training and runtime features: tree-structured drafting, parallel drafting, diffusion LLMs, DFlash, EAGLE.

  • Contributions to open-source projects and/or publications; please include links to GitHub pull requests, published papers and artifacts.

  • At NVIDIA, we believe artificial intelligence (AI) will fundamentally transform how people live and work. Our mission is to advance AI research and development to create groundbreaking technologies that enable anyone to harness the power of AI and benefit from its potential.

Our team consists of experts in AI, systems and performance optimization. Our leadership includes world-renowned experts in AI systems who have received multiple academic and industry research awards. If you’re excited to build systems, kernels, and tools that make large-scale AI faster, more efficient, and easier to deploy, we’d love to hear from you.


#LI-Hybrid

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 135,000 CAD - 185,000 CAD for Level 3, and 170,000 CAD - 220,000 CAD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 10, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

Similar roles

  • Gusto, Inc. logo

    Staff Software Engineer, Time and Scheduling

    Gusto, Inc.·Worldwide

    ---------------------------------------- About Gusto At Gusto, we're on a mission to grow the small business economy. We handle the hard stuff — payroll, health insurance, 401(k)s, and HR — so owners can focus on their craft and their customers. With teams in Denver, San Francisco, and New York, we support more than 500,000 small businesses nationwide and are building a workplace…

    • Remote
    • Full-time
  • Digital Zone logo

    Senior Site Reliability Engineer (Performance and Scalability)

    Digital Zone·United Arab Emirates

    Your mission is to make DigitalZone able to scale. You will build the platform's capacity to absorb campaign-level traffic spikes, and you will give every engineering team the tools, standards, and practices to load- and failure test their own systems. This is an enablement role at its core: you raise the reliability bar across the org by building capability, not…

    • Remote
    • Full-time
    • Green Visaself-sponsored, no tie
  • Squarepointcapital logo

    Software Developer - Risk Technology

    Squarepointcapital·London, United Kingdom

    <p><strong>Position Overview:</strong></p> <p>Risk Technology is a global team that designs, builds and maintains Squarepoint’s trading risk platform, which is responsible for trade capture, position management, profit/loss computation, inventory/locate management and internal order routing. These critical systems need to be performant, resilient, and capable of timely processing of high volumes of trading data in both live and historical scenarios, requiring solutions…

    • On-site
    • Full-time
  • Squarepointcapital logo

    Software Developer - Data Pipelines (Python)

    Squarepointcapital·London, United Kingdom

    <p><strong>Position Overview:</strong></p> <p>We are seeking an experienced Python developer to join our Alpha Data team, responsible for delivering a vast quantity of data served to users worldwide. You will be a cornerstone of a growing Data team, becoming a technical subject matter expert and developing strong working relationships with quant researchers, traders, and fellow colleagues across our Technology organization.</p> <p>Alpha…

    • On-site
    • Full-time
  • Squarepointcapital logo

    Junior Software Developer - Front-end

    Squarepointcapital·London, United Kingdom

    <p><strong>Please only apply to the one job you feel best fits your skillset and experience. If our team feels you are better suited for another role, we will reach out about the alternate opportunity.</strong></p> <p><strong>Position Overview:</strong></p> <p><span class="ui-provider a b c d e f g h i j k l m n o p q r s t u v…

    • On-site
    • Full-time
  • AoFrio logo

    Junior QA Engineer

    AoFrio·New Zealand

    Description Welcome to our World of Cold! At AoFrio, we are global leaders in providing IoT solutions to the food and beverage industry. Our innovative technology and dedicated team have positioned us at the forefront of our market. We are proud leaders in hardware-enabled Software as a Service (SaaS) for commercial refrigeration, with cutting-edge IoT solutions used by major brands…

    • Full-time
    • Skilled Migrant (Residence)PR, points-based, no single employer
  • AIA Group logo

    Software Engineer Specialist – Integration & AI

    AIA Group·New Zealand

    Your Role with Us As a Software Engineer Specialist – Integration in our Technology team, you’ll play a key role in designing and delivering integration solutions that keep our systems connected, secure, and running smoothly. You’ll collaborate with architects, engineers, and business stakeholders to build robust, scalable platforms that support both strategic initiatives and day-to-day operations. This is an exciting…

    • Full-time
    • Skilled Migrant (Residence)PR, points-based, no single employer
  • Tacto logo

    (Intern) AI Automation Engineer

    Tacto·Munich, Germany

    YOUR IMPACT As part of Solution Engineering at Tacto, you'll bridge our powerful platform with measurable customer value through data expertise. Working directly with customers and leads, you'll understand their unique supply chain challenges and implement tailored technical solutions. By integrating customer procurement data and configuring the platform to match their processes, you'll drive adoption and showcase immediate ROI (even…

    • On-site
    • Internship
    • Blue Cardtied; settle 21–33mo