Mirantis
Jobs at Mirantis
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Recently posted jobs
Software
Lead an engineering organization building secure enterprise container and Kubernetes platforms. Define technical strategy, improve platform resilience and operational excellence, embed security throughout the SDLC, balance feature delivery with stability, hire and develop engineering talent, and collaborate with product leaders, executives, and enterprise customers.
Software
Design and validate networking architectures for large-scale GPU clusters and AI infrastructure. Responsibilities include InfiniBand and RoCEv2 fabrics, multi-tenant networking, Linux host networking, Kubernetes and KubeVirt networking, automation, benchmarking, research, customer proofs of concept, technical documentation, partner collaboration, and conference presentations.
Software
Design and build Kubernetes-based LLM serving infrastructure, including GPU scheduling, scaling, model lifecycle management, Helm packaging, enterprise and offline deployments, API integrations, and production observability. Develop across a multi-service Go codebase, contribute design documentation and reviews, and help guide engineering direction. The role requires strong Kubernetes, distributed systems, CI/CD, and infrastructure-as-code expertise, with GPU, inference, or adjacent systems experience.
Software
Design and build Go-based storage control-plane services, CSI drivers, integrations, and tooling for Kubernetes AI and compute workloads. Automate storage provisioning using Terraform/OpenTofu and GitOps tools, support hybrid and bare-metal deployments, and develop observability for performance and reliability. Integrate enterprise storage platforms such as Dell PowerScale and VAST across cloud, edge, and air-gapped environments, including K0rdent, Cluster API, Harbor, and PKI/TLS infrastructure.
Software
Own end-to-end US recruiting for technical and business roles, including sourcing, pipeline development, hiring-manager advising, candidate experience, and closing. Recruiters will support engineering, product, account executive, pre-sales architect, and other go-to-market hiring. The role requires independent management of multiple searches, talent-market research, ATS administration in SmartRecruiters, and collaboration across distributed teams and time zones.
Software
Provides department-wide technical direction for k0rdent across Kubernetes, AI infrastructure, AI applications, and observability. Designs cross-cutting systems, establishes engineering standards, resolves architectural risks, mentors Staff and Principal engineers, and aligns product strategy with durable technical execution. The role also engages customers, partners, upstream communities, and industry forums while contributing code and prototypes to solve complex infrastructure problems.
Software
Build and own a multi-tenant web platform for managing AI infrastructure. Develop Next.js and React features, server-side authentication and API proxying, generated TypeScript SDKs, resilient data synchronization, design-system components, accessibility, performance, testing, observability, and production deployment. Collaborate with backend and globally distributed teams, review code, troubleshoot production issues, and mentor junior engineers.
Software
Design and implement infrastructure services for a GPU-as-a-Service platform. Build REST and gRPC APIs, bare-metal provisioning workflows, Kubernetes cluster automation, reconciliation loops, and durable asynchronous operations. Manage server enrollment, inspection, OS provisioning, cluster creation, hardware inventory interfaces, error handling, idempotency, and multi-tenant infrastructure reliability.
Software
Design and develop Rust and Go control-plane microservices, Kubernetes operators, reconciliation engines, and APIs for bare-metal AI infrastructure. Integrate datacenter networking systems, DPUs, network operating systems, and hardware-management platforms. Build observability with OpenTelemetry, troubleshoot distributed systems, and maintain rigorous testing practices. The role requires expertise in datacenter networking protocols, Linux networking, Kubernetes, provisioning, security, virtualization, and high-performance interconnects, along with architectural documentation and open-source collaboration.
Software
Design and build enterprise LLM-serving infrastructure on Kubernetes, including GPU scheduling, scaling, model lifecycle management, Helm packaging, offline deployments, platform integrations, and production observability. Contribute across a multi-service Go codebase, create design documentation, participate in reviews, and help guide engineering direction within an autonomous remote-first team.
Software
Own the vision, roadmap, backlog, and priorities for k0rdent AI observability across GPU infrastructure and distributed AI workloads. Define integrations for metrics, tracing, logs, telemetry, and alerting across Kubernetes, networking, storage, schedulers, inference platforms, and databases. Partner with engineering, marketing, field teams, customers, and ecosystem vendors to shape requirements, evaluate trade-offs, develop positioning, and support deployments.
Software
Deploys, integrates, operates, and tunes high-performance NFS storage for Kubernetes, GPU, and AI workloads. Responsibilities include CSI integration, Linux and network optimization, storage lifecycle management, air-gapped platform support, infrastructure-as-code, GitOps automation, observability, and troubleshooting distributed storage performance and reliability across hybrid, edge, and bare-metal environments.
Software
Deploy, integrate, operate, and optimize high-performance NFS storage for GPU-accelerated Kubernetes and AI platforms. Configure Linux storage and networking, integrate CSI storage into k0s and K0rdent clusters, automate provisioning through Terraform/OpenTofu and GitOps, and build observability for capacity, performance, and reliability. Diagnose end-to-end storage issues across hybrid, edge, and air-gapped environments while establishing operational standards.
Software
Design and build Go-based control-plane services, CSI drivers, storage integrations, observability tools, and infrastructure automation for Kubernetes AI platforms. Integrate enterprise and scale-out storage, support bare-metal and hybrid provisioning, and deliver declarative Terraform/OpenTofu and GitOps workflows. Architect reliable storage solutions for Kubernetes clusters, air-gapped environments, and sovereign infrastructure, including artifact mirrors and PKI/TLS.
Software
Own the networking vision and roadmap for k0rdent AI, defining underlay/overlay fabrics, RDMA, DNS/IPAM, network automation, and DPU/SmartNIC integration. Translate customer and engineering requirements into priorities, track emerging interconnect standards, and support positioning, field engagements, and reference architectures.
Software
Design, deploy, manage, and troubleshoot high-performance InfiniBand and Ethernet networks for HPC. Perform performance tuning, capacity planning, and monitoring. Implement Fortinet security, resolve routing/switching/latency issues, collaborate with compute and storage teams, document architectures, and participate in on-call escalation and upgrades.
One Month AgoSaved
Software
Lead operations for large-scale AI infrastructure: manage NVIDIA GPU and high-performance networking environments, troubleshoot Linux/Kubernetes/storage/hardware issues, drive incident response and root cause analysis, improve observability and automation, mentor engineers, and collaborate with engineering and datacenter teams to ensure platform reliability and scalability.
Software
Develop positioning and messaging for Mirantis AI infrastructure products; translate technical capabilities into audience-specific content; own competitive intelligence; create technical assets (white papers, briefs, reference architectures, sales enablement); support GTM, launches, analyst/media briefings, and industry representation.
Software
Operate and support production AI infrastructure across global datacenters, focusing on high-performance NVIDIA GPU platforms, Kubernetes, and high-speed networking. Monitor and troubleshoot infrastructure, network, hardware, and platform incidents; participate in incident response and root cause analysis; improve observability, automation, and operational runbooks.
