Senior Solutions Architect – AI Infrastructure Networking

  • Full-time

Company Description

About Mirantis

Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.  https://www.mirantis.com/

Job Description

About the role

We are seeking a Senior Solutions Architect with deep networking expertise to join the Voyager team in the Mirantis Office of the CTO. You will own the network solutioning of k0rdent AI, our platform for building and operating GPU clouds and AI factories, from metal to model.

You will design how large GPU clusters are interconnected, how tenants are isolated, and how workloads reach the network from containers, virtual machines and bare metal nodes. You will turn that design into published reference architectures, working code and proofs of concept that run on real hardware with our customers and partners.

As part of the Office of the CTO, research is a core part of the job. You will explore designs and technologies before customers ask for them, prototype ideas that may not ship, and challenge established practice when a better approach exists. Your work will shape where k0rdent AI goes next, not only how it is deployed today.

This role combines research with hands-on, outward-facing work. You will split your time between exploring new designs, writing architecture, building and validating it in the lab, and explaining it to engineers, executives and conference audiences. You will work with Solutions Architects, Partner Management and Engineering colleagues across many countries and time zones.

Key Responsibilities

Reference architecture

  • Design and publish network reference architectures and solution designs for k0rdent AI, covering data center topologies from a single rack to multi-thousand-GPU clusters.

  • Define compute (east-west) fabrics on InfiniBand and on RoCEv2 Ethernet, including rail-optimised and fat-tree / Clos designs, oversubscription and failure-domain choices.

  • Define front-end, storage and management networks, and how they connect to customer data centers, public clouds and hybrid environments.

  • Specify multi-tenant isolation across the stack: InfiniBand partitions (PKeys), VRFs and EVPN-VXLAN on Ethernet, and Kubernetes-level network policy.

  • Document scale limits and trade-offs (cost, performance, operability, vendor lock-in) with defensible reasoning a customer can follow.

 

Host and platform networking

  • Define how Linux hosts expose NICs, DPUs and SuperNICs to workloads: PCI passthrough, SR-IOV virtual functions, IOMMU groups, NUMA and GPU-NIC affinity.

  • Design Kubernetes networking for AI workloads: primary and secondary CNIs, Multus, SR-IOV and RDMA device plugins, NVIDIA Network Operator, and Dynamic Resource Allocation (DRA).

  • Design networking for virtual machines on KubeVirt, including passthrough and SR-IOV for GPU and RDMA traffic.

Research and exploration

  • Track and evaluate emerging technologies and standards in AI networking, such as Ultra Ethernet (UEC), SuperNIC and DPU offloads, scale-up and scale-across fabrics, and Kubernetes DRA network drivers.

  • Prototype alternative designs in the lab, including ones customers have not yet asked for, and measure them against current practice.

  • Challenge established or customer-preferred designs when evidence points to a better option, and make the case with data.

  • Publish findings as internal research notes and design proposals and, where appropriate, as external papers, blog posts or talks.

Code and proofs of concept

  • Build and run proofs of concept with customers and partners, on Mirantis lab hardware and on customer sites, and report results against agreed success criteria.

  • Write automation and tooling (Python, Go, Bash, Ansible, Helm, Kubernetes manifests, Terraform) that makes the reference architectures reproducible.

  • Run validation and benchmarking of fabrics and host configurations (e.g. NCCL tests, perftest, ib_write_bw) and turn findings into design guidance.

Customers, partners and community

  • Act as the network subject matter expert in customer discovery, design reviews and architecture workshops, bringing thought leadership and new ideas rather than only reflecting current practice.

  • Work with hardware and networking partners (e.g. NVIDIA, server OEMs, switch and network vendors) on joint designs and validations.

  • Innovate and feed research results back to Product and Engineering, and help shape the k0rdent AI roadmap.

  • Present at industry events, webinars and partner summits, and run hands-on workshops for technical audiences.

  • Write technical content: reference architecture documents, solution briefs, blog posts and internal enablement.

Qualifications

Education and experience

  • Bachelor's degree in Computer Science, Electrical Engineering, Telecommunications or a related field, or equivalent practical experience.

  • 8+ years in network engineering or network architecture, with at least 3 years in data center, HPC or cloud infrastructure networking.

  • Customer-facing experience as a solutions architect, pre-sales engineer, consultant or technical lead.

Required technical skills

  • Data center network design: spine-leaf and Clos topologies, rail-optimised GPU fabrics, oversubscription, ECMP, and the operational challenges of scaling to thousands of nodes.

  • InfiniBand: subnet management, partitioning, adaptive routing, and how NCCL / RDMA traffic behaves on the fabric.

  • RoCEv2 on Ethernet: lossless and congestion-controlled designs (PFC, ECN, DCQCN), QoS, MTU and buffer tuning.

  • Routing and overlays: BGP, EVPN-VXLAN, VRFs and their use for multi-tenancy.

  • Hybrid and multi-cloud connectivity: interconnects, VPN, cloud provider networking constructs, and IP address and DNS planning across sites.

  • Linux networking: how the kernel and drivers manage NICs, PCI passthrough, SR-IOV, IOMMU, NUMA affinity, and tools such as iproute2, ethtool and devlink.

  • Kubernetes networking: CNI plugins (e.g. Cilium, Calico, OVN-Kubernetes), Multus and secondary networks, SR-IOV and RDMA device plugins, network policy and service exposure.

  • Virtualisation on Kubernetes: KubeVirt, and how VM networking, passthrough and SR-IOV are configured there.

  • Programming or scripting in at least one language (Python or Go preferred), and comfortable using Git, CI and infrastructure-as-code.

How you work

We hire for these behaviours as much as for technical depth, and will ask for concrete examples of each.

  • Curious. You dig into why a fabric behaves the way it does, read the specs and the source, and keep up with how AI networking is changing. You question established designs, including the ones customers are used to, when there is a better way.

  • Ownership. You take a reference architecture or a POC from the first sketch to a published, validated result, and you stand behind it.

  • Proactive. You spot gaps in the product, the documentation or a customer design before you are asked, and you act on them.

  • Fast learner. You get productive quickly in unfamiliar hardware, software or customer environments.

  • Business acumen. You connect network design choices to cost, time-to-value and risk, and explain them in those terms to decision makers.

Communication and collaboration

  • Excellent written and spoken English; you write documents others can build from.

  • Comfortable presenting to large audiences at conferences and events, and running hands-on workshops.

  • Able to adapt the message to the audience, from network engineers to customer executives.

  • Experience working in an international, distributed company, across time zones and different work cultures, with Solutions Architects, Partner Management and Engineering teams.

Nice to have

  • Hands-on experience with NVIDIA networking: Quantum InfiniBand, Spectrum-X Ethernet, ConnectX SuperNICs, BlueField DPUs, UFM, and NVIDIA reference designs for AI factories.

  • Experience with GPU cloud, neocloud or HPC operators, or with NVIDIA Cloud Partner (NCP) style deployments.

  • Bare-metal provisioning and lifecycle: Metal3, Ironic, Redfish, PXE / iPXE, and Cluster API.

  • Network automation and SDN controllers, and switch operating systems such as Cumulus Linux, SONiC or vendor equivalents.

  • Global load balancing and service exposure for distributed inference (GSLB, Gateway API, ingress).

  • Familiarity with Software-Defined Storage networking for AI (NVMe-oF, GPUDirect Storage).

  • Experience with Mirantis products, MKE, k0s, k0rdent networking.

  • Contributions to open-source networking or Kubernetes projects, or a track record of public talks and technical writing.

  • Participation in standards or industry bodies (e.g. Ultra Ethernet Consortium, OCP, IETF), or published research, patents or white papers.

  • Additional languages beyond English, for example a European language.

Location, travel and working model

This role is remote, open to candidates based in Europe, the United States (East Coast preferred) or India.

  • You will work with a team spread across many time zones; some meetings will fall outside standard local hours.

  • Travel of up to 25% for customer engagements, partner meetings, lab work and industry events, primarily within the EU and US.

  • Mirantis is an equal opportunity employer. Compensation and benefits are set according to the local market and employment arrangement.

Additional Information

Why you’ll love Mirantis

  • Build the observability foundation for the AI cloud era, working directly with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises

  • Collaborate with a world-class, distributed team committed to openness and technical excellence

  • Shape the product narrative and influence go-to-market success

 

It is understood that Mirantis, Inc. may use automated decision-making technology (ADMT) for specific employment-related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to [email protected]

By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for current and future job opportunities.

#remote

We are a Leader for Container Management in G2 (#2 after AWS)!

By clicking the link above or any third-party link within this posting, you are leaving this site and going to a third-party website where the third-party website's terms and privacy policy apply

Privacy Notice