xAI Logo

xAI

Lead, Hardware Deployment Engineer

Posted Yesterday
Be an Early Applicant
In-Office or Remote
Hiring Remotely in Memphis, TN
Senior level
In-Office or Remote
Hiring Remotely in Memphis, TN
Senior level
Lead and build an in-house hardware deployment team to own L11 rack integration, GPU hardware bring-up, post-L11 repair, vendor SLA enforcement, root-cause analysis, and process/tooling to maximize node availability across multiple data halls.
The summary above was generated by AI

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

As the Hardware Deployment Engineer Lead, you will own the end-to-end bring-up of GPU compute hardware across the world's largest AI training clusters. You will build and lead a dedicated in-house hardware deployment team responsible for L11 integration, hardware bring-up, and post-L11 repair of GB300-class systems across multiple data halls concurrently. Your team's throughput directly determines how fast xAI can compute online — this is one of the most critical path activities in the company. You will set the deployment playbook, hold hardware vendors accountable to SLAs, and institutionalize processes so that cluster deployment is limited only by hardware supply and power, never by deployment velocity. The position is based in Memphis, TN.

RESPONSIBILITIES:
  • Lead, hire, and develop a dedicated hardware deployment team (deployment engineers, deployment technicians, and repair technicians) with full ownership of team structure and staffing.
  • Own L11 rack integration and compute hardware bring-up across multiple data halls concurrently, from delivery dock to healthy production handoff.
  • Drive aggressive bring-up timelines: achieve 95%+ node availability within days of rack delivery and 100% closure within one week per data hall.
  • Own post-L11 hardware health: run systematic health pushes to sustain greater than 98% node availability prior to turnover to operations.
  • Internalize non-RMA hardware repairs to maximize hardware recovery, minimize repair backlogs, and reduce dependence on OEM turnaround times.
  • Develop and enforce vendor SLAs for OEM and supplier responsibilities; prevent accumulation of unrepaired hardware ("bone piles") and repair backlogs before turnover to operations.
  • Perform root cause analysis of hardware failures discovered during L11 and drive corrective actions with vendors and internal engineering teams.
  • Partner with site operations on hardware debugging and repair, and train site operations teams to support future data center deployments.
  • Build, document, and continuously improve deployment processes, tooling, and training so bring-up capability scales across sites and future hardware generations.
BASIC QUALIFICATIONS:
  • 5+ years of hands-on experience deploying, integrating, or repairing compute/server hardware at data center scale.
  • Direct experience with L11 (rack-level) integration and bring-up of GPU or accelerator-based systems.
  • Demonstrated experience leading technician or engineering teams in a fast-paced deployment, manufacturing, or data center environment.
  • Deep troubleshooting skills across servers, GPUs, NVLink/fabric interconnects, high-speed networking, and liquid cooling systems.
  • Willingness to work on-site in Memphis, TN, including extended hours and weekends during critical bring-up phases.
PREFERRED SKILLS AND EXPERIENCE:
  • Experience with NVIDIA GB200/GB300 NVL72 or similar rack-scale liquid-cooled GPU systems.
  • Experience standing up a new team or function, including hiring, training, and process development from scratch.
  • Experience managing OEM/ODM vendor relationships (e.g., Dell, Supermicro), including SLA definition and enforcement.
  • Experience with hardware failure analysis, RMA processes, and component-level repair strategies at fleet scale.
  • Experience with data center automation, burn-in/validation tooling, and hardware health telemetry.
  • Track record of driving step-change improvements in deployment velocity or cost (e.g., insourcing work previously performed by OEMs).

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.

Similar Jobs

27 Minutes Ago
Easy Apply
Remote or Hybrid
USA
Easy Apply
55K-68K Annually
Junior
55K-68K Annually
Junior
Fintech • Information Technology • Payments • Productivity • Software • Travel • Automation
Provide end-to-end venue sourcing and high-touch client liaison for events across the US. Build client and internal partnerships, negotiate venue terms, maximize revenue and ROI, maintain project documentation, generate leads, and continually update product knowledge through site visits and training.
Top Skills: CventGoogle WorkspaceMS OfficeNavan Systems
28 Minutes Ago
Easy Apply
Remote or Hybrid
USA
Easy Apply
Junior
Junior
Fintech • Information Technology • Payments • Productivity • Software • Travel • Automation
Support Event Planners as primary client liaison for multi-element, mid and large projects. Manage small meetings (10–15 guests), negotiate with venues, track contracts and commissions, assist with PNRs/WEX cards, maintain documentation, take meeting minutes, support invoicing/reconciliation, escalate upsell opportunities, and provide occasional onsite/travel support. Use Cvent and internal systems for delegate management.
Top Skills: CventGoogle SuiteMS Office
4 Hours Ago
Remote or Hybrid
United States
180K-215K Annually
Senior level
180K-215K Annually
Senior level
HR Tech • Information Technology • Professional Services • Sales • Software
Lead U.S. product efforts for Time Off & Attendance, owning vision, strategy, roadmap, and execution. Partner cross-functionally with Engineering, Design, Payroll, Sales, and Customer Success to build compliant, scalable workforce management and payroll integrations. Use customer research and data to define metrics, iterate, and embed AI capabilities. Serve as founding U.S. PM and work closely with global teams.
Top Skills: AI

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account