Nvidia brand banner

Senior Software Engineer, Automation Infrastructure

NvidiaShanghai, ChinaPosted 5d ago
via Workday

NVIDIA's Performance Lab (PerfLab) builds the systems and automation used to evaluate the performance and quality of accelerated computing and AI workloads. We turn complex benchmark experiments into reliable, scalable, and reproducible workflows that help engineering teams make better decisions faster.

We are looking for an experienced and highly self-motivated System Software Engineer to help build the next generation of our infrastructure. You will independently own meaningful platform components and lead projects from problem discovery and technical design through production deployment and adoption. The ideal candidate enjoys finding important engineering problems, understanding their root causes, and using technology to create simple, reusable solutions.

You will collaborate with NVIDIA teams around the world and influence how we evaluate evolving areas such as large language models, agentic AI, accelerated computing, and other emerging AI workloads.

What You'll Be Doing

  • Define the technical direction and architecture for major areas of PerfLab's benchmark infrastructure, translating evolving business and engineering needs into clear roadmaps and scalable platform capabilities.

  • Lead the design and implementation of reusable software, services, and workflows that automate benchmark definition, execution, result collection, validation, and reporting across local, cluster, and cloud-native environments.

  • Remain hands-on with performance testing and analysis, developing a deep understanding of existing workflows and using that knowledge to guide platform investments and technical decisions.

  • Establish engineering approaches that improve the reliability, scalability, observability, maintainability, and reproducibility of large benchmark campaigns.

  • Lead the diagnosis of complex, cross-layer issues spanning applications, Linux systems, containers, distributed jobs, compute resources, networking, and storage.

  • Build strong partnerships with performance engineers, QA teams, product teams, and other customers; create alignment across organizations and drive high-impact ideas and projects from concept through adoption.

  • Provide technical leadership through architecture and code reviews, clear decision-making, high engineering standards, and mentoring of other engineers.

  • Identify emerging technologies, including AI-assisted automation, and determine where they can deliver meaningful improvements in benchmark creation, failure triage, data analysis, or engineering productivity.

  • Contribute to the strategy and development of internal and open-source infrastructure projects, and help grow their adoption across teams.

What We Need to See

  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent practical experience.

  • 6+ years of relevant software engineering experience in system software, infrastructure, developer platforms, distributed systems, or production automation.

  • Strong Python programming and software engineering expertise, with a track record of building and operating production-quality tools, services, or automation frameworks.

  • Solid understanding of Linux and system-level concepts such as processes, concurrency, networking, storage, resource management, and failure handling.

  • Extensive experience with containers and workload orchestration, scheduling, or distributed computing platforms, including the design of reliable systems with clear interfaces, testing, observability, and recovery behavior.

  • Knowledge of machine learning, AI, or accelerated-computing workloads and experience reasoning about their performance, quality, and operational tradeoffs.

  • Proven technical leadership on sophisticated, multi-functional projects, including defining architecture, managing technical risk, resolving ambiguity, and driving solutions through delivery and adoption.

  • Strong analytical and problem-solving abilities, with the judgment to prioritize effectively, manage multiple initiatives, and adapt in a dynamic, constantly evolving environment.

  • Excellent communication, organizational, and influencing skills, with the ability to align globally distributed collaborators and drive decisions.

  • A record of mentoring engineers, elevating engineering quality, and helping teams make better technical decisions.

Ways to Stand Out From the Crowd

  • Experience leading major initiatives in GPU or AI infrastructure, distributed training or inference, model evaluation, or performance benchmarking.

  • Experience architecting workflow engines, schedulers, experiment platforms, test frameworks, or developer infrastructure used by multiple teams.

  • Experience operating distributed or cloud-native systems at scale, including performance profiling, capacity analysis, resource scheduling, or multi-node workloads.

  • Practical experience establishing AI-agent, tool-calling, or coding-agent strategies that improved engineering workflows at team or organizational scale.

  • Demonstrated success turning loosely defined, cross-organizational problems into durable platforms or programs with measurable engineering impact.

We have some of the most forward-thinking and hardworking people in the world working for us. If you are creative, autonomous, and passionate about providing technical leadership while remaining hands-on in building systems that make complex AI performance work repeatable and scalable, we want to hear from you.

Similar roles

  • AIA Group logo

    Software Engineer Specialist – Integration & AI

    AIA Group·New Zealand

    Your Role with Us As a Software Engineer Specialist – Integration in our Technology team, you’ll play a key role in designing and delivering integration solutions that keep our systems connected, secure, and running smoothly. You’ll collaborate with architects, engineers, and business stakeholders to build robust, scalable platforms that support both strategic initiatives and day-to-day operations. This is an exciting…

    • Full-time
    • Skilled Migrant (Residence)PR, points-based, no single employer
  • Sunstone Talent logo

    Senior Web Developer (Shopify & Webflow?)

    Sunstone Talent·Christchurch, New Zealand

    An exciting sustainable green tech company is looking for a Senior Web Developer who has passion developing & delivering new features for leading edge products? Come join a great team with purpose? What you’ll bring: * BSc, BEng or BA or similar or not * 4 years+ of web development experience including Webflow & Shopify building & enhancing eCommerce websites…

    • Full-time
    • NZD 110k–NZD 130k / yr
    • Skilled Migrant (Residence)PR, points-based, no single employer
  • VISPIRON GmbH logo

    Senior Software Engineer (m/w/d) - SEtrade

    VISPIRON GmbH·Munich, Germany

    Strom lässt sich nicht auf Halde legen. Erzeugung und Verbrauch müssen sich in jeder Sekunde ausgleichen, sonst kippt die Netzfrequenz. Das ist keine Haltung, das ist Physik. Aus dieser Physik entsteht ein Markt, der in 15-Minuten-Scheiben handelt und keine Rücksicht nimmt. Genau dort läuft unsere Software. Wir bewirtschaften heute 350 MWp Solar, 100 MWp Wind, 100 MWh Speicher, 12 MW…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • Addepar1 logo

    Engineering Manager - Data Platform

    Addepar1·Edinburgh, United Kingdom

    <div class="content-intro"><p><span style="text-decoration: underline;"><strong>Who We Are</strong></span></p> <p>Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. Addepar unifies portfolio, market and client data in a total portfolio view and delivers AI-powered insights within investment and client workflows. More than 1,400 firms in nearly 60 countries use Addepar to manage and advise…

    • On-site
    • Full-time
  • Checkout.com logo

    Senior Mobile Engineer

    Checkout.com·London, Canada

    Company Description We’re Checkout.com. You might not know our name, but companies like eBay, Spotify, Klarna, Uber, and Sony do, because we’re behind many of the digital experiences you use every day. We are where the world checks out, enabling over 10 billion transactions yearly for more than one billion global shoppers. Whether you want to book a holiday, order…

    • On-site
    • Full-time
    • Express EntryPR day one, no employer
  • Stark logo

    Software Integration & Test Engineer (Maritime Autonomy) (all genders)

    Stark·Plymouth, United Kingdom

    About Us STARK is a new kind of defence technology company revolutionizing the way autonomous systems are deployed across multiple domains. We design, develop, and manufacture high-performance unmanned systems that are software-defined, mass-scalable, and cost-effective. This provides our operators with a decisive edge in highly contested environments. We’re focused on delivering deployable, high-performance systems—not future promises. In a time of…

    • On-site
    • Full-time
  • Clera logo

    Platform Engineer

    Clera·Berlin, Germany

    ABOUT THE ROLE A fast-growing, Berlin-based enterprise AI platform startup is looking for a Platform Engineer to help build and maintain the backbone of their multi-tenant SaaS infrastructure. Founded in 2023 and backed by notable investors, the company serves thousands of enterprise customers and is scaling rapidly. You'll be a critical hire on a small, high-calibre team, working on-site in…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo
  • Clera logo

    Full-Stack Software Engineer

    Clera·Berlin, Germany

    ABOUT THE ROLE A fast-growing, seed-stage enterprise AI platform company based in Berlin is looking for a Full-Stack Software Engineer to join their small, high-impact engineering team. Founded in 2023, the company builds an all-in-one platform that enables organisations to securely adopt and integrate generative AI — including AI chat, workflows, and custom agents — across their entire operations. This…

    • On-site
    • Full-time
    • Blue Cardtied; settle 21–33mo