Radiant brand banner

Cluster Architect

RadiantLondon, CanadaPosted 3h agoonsite
via Arbeitnow

About Radiant

Radiant is redefining how AI infrastructure is built.

We design and operate AI-native cloud platforms engineered for sovereignty, performance, and scale. Our infrastructure powers GPU-native workloads, multi-tenant control planes, and high-performance AI systems designed for the most demanding environments.

We are not building a generic cloud. We are building purpose-built AI infrastructure - from powered land, to compute, to software .

As we scale our platform and expand our engineering organisation, we are looking for leaders who can build strong teams, uphold high standards, and deliver reliably at pace.

As part of the Infrastructure Solutions team, the Cluster Architect is the tip of the spear leading Radiant’s new GPU deployments. The Cluster Architect contributes to the technical strategy, reference architecture, and governance standards underpinning Radiant's GPU compute platform. They translate the high-level designs provided by the Customer Success organization into fully operational supercomputers. In conjunction with the Datacenter Engineering team, they lead the definition of the low-level design of upcoming high-performance computing (HPC) and artificial intelligence (AI) infrastructure deployments.

You must be a solution finder, able to adapt to evolving technology and customer requirements, and be ready to handle scale. The right person for this role will be technically focused and detail-oriented, and able to work with the broader team to successfully deliver some of the world’s largest supercomputers at a rapid pace with precision. It is critical to document each step of the process to enable effective communication with internal and external stakeholders.

About the Infrastructure Team

The Infrastructure Engineering team spearheads the strategy, roadmap, and delivery of critical HPC/AI infrastructure to the broader Technology and Customer organizations. Its mission is to deploy scalable, secure, and efficient infrastructure solutions tailored to the organization's strategic needs. The team defines Radiant’s supercomputers’ reference architecture in line with the customer, software, and operational requirements of the company.

They work closely with third-party suppliers to provide and maintain the list of technologies that form part of Radiant’s portfolio of offerings.

They provide technical support throughout the procurement process with vendors to ensure the suitability of RFI/RFP responses.


They work in tandem with OEMs and infrastructure providers during the infrastructure rollout, and providing technical inputs to DC providers and OEMs

They work closely with the program management office, and align the hardware infrastructure rollout programs with the parallel datacenter fitout and software engineering programs.

They play an instrumental role in the successful deployments of cloud infrastructure, whilst ensuring optimal performance and reliability for high-demand computing environments. Once accepted, the infrastructure solutions are then handed over to the Infrastructure Operations team who manages the daily management of the infrastructure.

What you’ll do

  • Engage with key stakeholders (such as the Program Management Office, Data Center Engineering, Solutions Engineering and Infrastructure Operations teams) and translate inputs into Radiant’s reference architecture and low-level designs blueprints.

  • Remain abreast of technology advancements and roadmaps with key vendors in the space.

  • Lead the low-level design of high-density AI supercomputer systems (compute, interconnect, storage, management).

  • Participate in the Architecture Review Board involving key stakeholders to ensure collaboration and company alignment behind the designs being produced.

  • Evaluate tradeoffs between performance, power, cooling, serviceability, and cost throughout the procurement process.

  • In tandem with potential suppliers, drive the creation of detailed low-level designs, including (but not limited to):

    • Rack elevation diagrams and layouts

    • Data center floor plans and spatial planning

    • Power distribution and cooling architecture and related detailed specifications

    • Detailed Bill of Materials (BoM)

    • Port mapping documents.

  • Drive thorough technical evaluations and response accuracy from vendors during the RFx processes, ensuring designs are aligned with vendor capabilities and long-lead item constraints.

  • Play a hands-on and pro-active role in the manufacturing, burn-in, cabling, logical configuration, benchmarking, and logical implementation of new deployments to turn designs into production-grade systems.

  • Drive acceptance criteria with quality, reliability, safety, and performance goals in mind.

  • Work with the Software Engineering and Solutions Engineering teams to ensure the infrastructure supports AI workload requirements.

  • Engage with project management to define milestones, risks, resourcing, and deliverables associated with the delivery.

  • Produce accurate architecture diagrams, configuration documentation, runbooks, and operational guidance to facilitate the handover of the new deployments to the Infrastructure Operations team.

  • Establish delivery best practices and continuously improve the designs, processes, and documentation through the participation of “lessons learned” workshops.

Requirements

  • BSc in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or other related Engineering fields.

  • 5+ years of solution architecture experience, demonstrating motivation and skills to drive technical design and delivery processes.

  • Deep expertise in datacenter engineering (mechanical, electrical and plumbing), GPU systems, high performance networking (like InfiniBand), including a proven understanding of network topologies and storage architecture.

  • Proficiency in system-level aspects, encompassing Operating Systems, Linux kernel drivers, GPUs, NICs, and infrastructure software solutions.

  • Demonstrated experience producing detailed infrastructure designs (including network topologies, rack layouts, BoMs, power/cooling planning…)

  • Deep understanding in datacenter fit-out, and power/cooling constraints

  • Excellent interpersonal and communication skills, with a track record interfacing with various types of stakeholders (engineering/operations teams, vendors, etc.) and writing clear deliverables.

  • Comfortable working in fast-moving, high-complexity technical environments.

Nice to have

  • Prior work experience in HPC, Telco, or hyperscale supercomputing environments designing, delivering, and/or operating large scale HPC systems.

  • Work or research experience in networking fundamentals, TCP/IP stack, and data center designs.

  • Familiarity with monitoring/telemetry systems and observability for infrastructure health.

  • Demonstrated expertise in cloud orchestration software and job schedulers, including platforms like Kubernetes, Docker Swarm, and HPC-specific schedulers such as Slurm.

  • Knowledge of DevOps/MLOps technologies such as Docker/containers, Kubernetes, datacenter compute/network/storage deployments.

  • Hands-on experience with NVIDIA systems/SDKs (e.g., CUDA), NVIDIA Networking technologies (e.g., DPU, RoCE, InfiniBand), ARM CPU solutions, coupled with proficiency in C/C++ programming, parallel programming, and GPU development.

 

Why should you join us?

 

What sets us apart is our blend of modern technology, competitive benefits, and an open, welcoming work culture that enables our people to thrive.

 

Here are just some of the great things you can expect from us:

  • 25 days of annual leave

  • A culture that emphasises results over hierarchy, process & ego: we place great emphasis on the quality, ingenuity and creativity of work.

  • Open communication, regular feedback: we value smooth collaboration, direct and actionable feedback, and believe that leading with empathy and a growth mindset makes us better together.

  • Learning Time: we all have dedicated learning time to focus on new skills, projects or interests that lay outside of your day-to-day job.

  • Health & Wellbeing: we want everyone to feel healthy and happy, so we offer private medical insurance via Bupa.

  • Cycle to Work Scheme: we're committed to building a sustainable business, so we encourage cycling to work.

  • Gympass subscription to a variety of gyms and wellbeing apps

  • Participation in the company shares program

  • Enhanced parental pay & leave

Diversity, Equality, Inclusion and Belonging

We are an equal opportunity employer and we strive to reduce unconscious bias throughout our hiring process. All applicants will be considered for employment without attention to ethnicity, religion, sexual orientation, gender identity, family or parental status, national origin, veteran, neurodiversity status or disability status. To ensure our recruitment processes provide an equal opportunity for all applicants to succeed, we encourage you to let us know if there are any adjustments that we can make.

Find more English Speaking Jobs in United Kingdom on Arbeitnow

Skills

  • Infrastructure Engineering

Similar roles

  • A2MAC1 logo

    Senior Costing Engineer ( Chassis, Thermal, Battery)

    A2MAC1·India

    Position Value 1.Understand latest and cost-efficient engineering solution through technical benchmark 2.Learn the engineering excellent practices from high-performance vehicle products across the world 3.Co-work with OEM or Tiers to attend in new product development phase 4.Communicate with global technical or costing experts 5.Utilize technical insights database Responsibilities and duties 1.Complete and report relevant analysis reports according to customer requirements 2.Research…

    • Full-time
  • A2MAC1 logo

    Costing Engineer ( Seat , Soft Trims)

    A2MAC1·India

    Position Value 1.Understand latest and cost-efficient engineering solution through technical benchmark 2.Learn the engineering excellent practices from high-performance vehicle products across the world 3.Co-work with OEM or Tiers to attend in new product development phase 4.Communicate with global technical or costing experts 5.Utilize technical insights database Responsibilities and duties 1.Complete and report relevant analysis reports according to customer requirements 2.Research…

    • Full-time
  • Aviso Wealth logo

    System Support Technician I

    Aviso Wealth·Vancouver, Canada

    Aviso Wealth: At Aviso, we are dedicated to improving the financial well-being of Canadians. As a leading wealth management organization, we are committed to leadership, innovation, partnership, responsibility, and community. Working with talented and energetic professionals who exemplify our values every day, you will quickly notice that our people and dynamic ‘oneaviso’ culture sets us apart. If you are looking…

    • Hybrid
    • Full-time
    • Express EntryPR day one, no employer
  • Gocardless logo

    Senior Global Performance Marketing Manager

    Gocardless·London, Canada

    <div class="content-intro"><p><strong>About us</strong></p> <p>GoCardless is a <strong>global bank payment</strong> company. Over <strong>100,000 businesses</strong>, from start-ups to household names, use GoCardless to collect and send payments through direct debit, real-time payments and open banking.&nbsp;</p> <p>GoCardless processes <strong>US$130bn+</strong> of payments annually, across <strong>30+ countries</strong>; helping customers collect and send both <strong>recurring</strong> and <strong>one-off payments</strong>, without the chasing, stress or expensive fees. We…

    • On-site
    • Full-time
    • Express EntryPR day one, no employer
  • rtbhouse logo

    Account Executive

    rtbhouse·London, Canada

    <p><strong>Role: </strong>Account Executive</p> <p><strong>Location: </strong>United Kingdom</p> <p><strong><span style="color: rgb(224, 62, 45);">We Are</span>:&nbsp;</strong></p> <p>RTB House is a next-generation performance demand-side platform (DSP) that uses proprietary Deep Learning AI algorithms to help brands grow. The company is the market leader in driving performance using Deep Learning across the entire purchase funnel.</p> <p>Founded in 2012 and now operating in 90+ markets, RTB House…

    • On-site
    • Full-time
    • Express EntryPR day one, no employer
  • speechify logo

    Go-to-Market - London, United Kingdom

    speechify·London, Canada

    <p><strong>Mission</strong></p> <p>The mission of Speechify is to make sure that reading is never a barrier to learning. Over 50 million people use Speechify's text-to-speech products to turn whatever they're reading – PDFs, books, Google Docs, news articles, websites – into audio, so they can read faster, read more, and remember more. Google recently named Speechify the Chrome Extension of the…

    • On-site
    • Full-time
    • Express EntryPR day one, no employer
  • Obsidiansecurity logo

    Enterprise Account Executive - UK

    Obsidiansecurity·London, Canada

    <div class="content-intro"><p>Obsidian Security is the leading SaaS security platform, trusted by global enterprises like Snowflake, T-Mobile, and Algolia. We protect 200+ organizations across North America, Europe, the Middle East, Southeast Asia, Australia, and New Zealand, including many of the world’s largest Fortune 1000 and Global 2000 companies.</p> <p>Founded in 2017 and backed by top investors like Greylock, Obsidian was built…

    • On-site
    • Full-time
    • Express EntryPR day one, no employer
  • Checkout.com logo

    Senior Director, Reward, Performance and Analytics

    Checkout.com·London, Canada

    Company Description We’re Checkout.com. You might not know our name, but companies like eBay, Spotify, Klarna, Uber, and Sony do, because we’re behind many of the digital experiences you use every day. We are where the world checks out, enabling over 10 billion transactions yearly for more than one billion global shoppers. Whether you want to book a holiday, order…

    • On-site
    • Full-time
    • Express EntryPR day one, no employer