
Software Engineer, LLM Inference
NVIDIA has continuously reinvented itself over two decades. NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world.
This is our life’s work — to amplify human imagination and intelligence. AI becomes more and more important in AI-City and self-driving car. NVIDIA is at the forefront of the AI-City and self-driving revolution and providing powerful solutions for them. All these solutions are based on GPU-accelerated libraries, such as CUDA, cuDNN and TensorRT, etc. Now, we are now looking for an CPU computing engineer based in Shanghai.
What you’ll be doing:
Craft and develop robust inferencing software that can be scaled to multiple platforms for functionality and performance
Performance analysis, optimization and tuning
Closely follow academic developments in the field of artificial intelligence and feature update TensorRT and TensorRT Edge LLM
Collaborate across the company to guide the direction of machine learning inferencing, working with software, research and product teams
What we need to see:
Masters or higher degree in Computer Engineering, Computer Science, Applied Mathematics or related computing focused degree (or equivalent experience)
4+ years of relevant software development experience.
Excellent C/C++ programming and software design skills, including debugging, performance analysis, and test design.
Strong curiosity about artificial intelligence, awareness of the latest developments in deep learning like LLMs, generative models
Experience working with deep learning frameworks like PyTorch
Proactive and able to work without supervision
Excellent written and oral communication skills in English
Strong customer communication skills, powerfully motivated to provide highly responsive support as needed
Similar roles
Software Manager, Robotics Platform Engineering
Nvidia·Shanghai, China
NVIDIA has transformed computer graphics, PC gaming, and accelerated computing for more than 25 years through groundbreaking technology and exceptional people. Today, we are harnessing the potential of AI to define the next era of computing, where GPUs power computers, robots, and autonomous vehicles that can understand and interact with the world. We are seeking a Software Manager to lead…
- Full-time
Senior Software QA Test Developer - Embedded
Nvidia·Pune, India
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by phenomenal technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving…
- Full-time
Senior Software Engineer - Husky Storage
Datadog·Boston, United States
Husky is what we call the distributed, petabyte-scale columnar event store at the heart of our Event Platform, which powers dozens of Datadog’s most popular products – Logs, RUM, APM, Cloud Network Monitoring, Netflow, and many more. Husky was built from the ground up at Datadog to store and query massive volumes of event data at low cost, with data…
- Full-time
Full-Stack Software Engineer, Emerging Products
Openai·San Francisco, United States
About the Team The Emerging Products team is a lean, high-output product lab group that builds products at the forefront of model capabilities. We collaborate across all teams within the company, from research and infrastructure to consumer products. The team is responsible for identifying new product opportunities, building them quickly, dogfooding them internally, and then launching the successful products to…
- Full-time
Full-Stack Engineer - Creative Studio
Elevenlabs·United Kingdom
ABOUT ELEVENLABS ElevenLabs is an AI research and product company transforming how we interact with technology. We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses - from fast-growing startups to large enterprises like Deutsche Telekom and Meta. Our investors are some of the world's most prominent, including…
- On-site
- Full-time
Backend Engineer, IAM
Reddit·United States
Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s…
- Full-time
Engineering Manager, International
Robinhood·Toronto, Canada
JOIN US IN BUILDING THE FUTURE OF FINANCE. Our mission is to democratize finance for all. An estimated $124 trillion of assets will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. ABOUT THE…
- Full-time
- Express Entry — PR day one, no employer
Senior Salesforce Developer (Temporary)
Angi·United Kingdom
For over 30 years, Angi has powered the future of the home services industry, creating an environment where homeowners and pros benefit from more jobs done well. For homeowners, our platform is a reliable way to find skilled pros. For pros, we're a reliable business partner who helps them find the winnable work they want, when they want. For employees,…
- On-site
- Full-time