Fuse Energy Logo

Fuse Energy

AI Inference Engineer

Posted 21 Days Ago
In-Office or Remote
Hiring Remotely in United States
Mid level
In-Office or Remote
Hiring Remotely in United States
Mid level
Define and build Fuse's inference-serving architecture and stack, including request routing, batching, scheduling, and autoscaling for latency-sensitive workloads. Own model-level optimizations (quantisation, distillation, speculative decoding), translate performance SLAs into capacity plans, and integrate low-level GPU/CUDA performance work with serving infrastructure. Set benchmarks, tooling, and act as technical owner of inference performance and reliability.
The summary above was generated by AI

Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy - fast. We're combining first-principles thinking with cutting-edge technology to build a radically better energy system. We raised $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group and strategic angels like Nico Rosberg, the Co-Founder of Solana and GPs behind Meta, Revolut, Spotify, Uber and more.
As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure that sits at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch - and we're looking for the founding engineer to own the latter.
We're looking for a Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, reporting directly to the CTO. Where our CUDA and GPU engineering hires own kernel-level and hardware performance, this role owns the layer above it: how models actually get served, scaled, and delivered against committed performance targets.

The Opportunity

Fuse is seeing significant demand for data centre capacity across the markets we operate in, primarily for inference. Few companies in the world can pair real power delivery with real compute the way Fuse can, which puts inference serving at the heart of how we turn that advantage into the best offering in the market. That's this role.

Responsibilities

    • Define Fuse's inference serving strategy and architecture from first principles.
    • Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads.
    • Own model-level optimisation strategy for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers.
    • Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents).
    • Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans.
    • Act as a direct technical owner of inference performance and reliability.
    • Work closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer.
    • Set the standards, tooling, and benchmarks this function will run on as it grows.

Requirements
  • 4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience.
  • Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding).
  • Strong systems thinking - able to reason about the full path from incoming request to served response across a large cluster.
  • Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system.
  • A track record of making high-stakes architecture calls and owning the outcome.
  • Comfort operating without a playbook - this is a founding role shaping a new function around architecture that's still early-stage, not joining an established one.

Nice to Have

  • Experience with Triton or custom ML inference/training frameworks.
  • Experience with autoscaling or capacity planning for large-scale inference workloads.
  • Exposure to multi-tenant serving or SLA-driven infrastructure.
  • Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system.
  • Familiarity with Kubernetes/Slurm for cluster orchestration.
  • Interest or experience in energy markets, grid systems, or sustainability-focused compute.

Benefits
  • Competitive salary and an equity sign-on bonus.
  • Biannual bonus scheme.
  • Fully expensed tech to match your needs.
  • Breakfast and dinner allowance for office based employees.
HQ

Fuse Energy Greenwood Village, Colorado, USA Office

Greenwood Village, CO, United States

Similar Jobs

Senior level
Blockchain • Software • Analytics • Financial Services • Cryptocurrency
As an AI Research Engineer, you will optimize model serving and inference architectures, focusing on high-performance, resource-efficient solutions for AI systems across various deployment environments.
Top Skills: AIDiffusion ModelsExpert ParallelismFlash AttentionGpu KernelsInference OptimizationKernel OptimizationKv CacheMetal Shading LanguageMobile DevicesPipeline ParallelismPruningQuantizationSpeculative DecodingTensor ParallelismVision Transformers
14 Days Ago
In-Office or Remote
Mid level
Mid level
Blockchain • Software • Analytics • Financial Services • Cryptocurrency
As an AI Inference Engineer, you will optimize C++ systems for AI model inference on edge devices, collaborating with researchers to deploy and enhance models while ensuring runtime stability and performance.
Top Skills: C++GgmlLlama.CppOnnx
An Hour Ago
Remote
Senior level
Senior level
Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
Serve as the L2 SME for Netcracker Convergent Billing & Rating, handling P1/P2 incidents, RCA, troubleshooting billing/rating/charging issues, validating data consistency across catalog and engines, coordinating escalations and releases, maintaining runbooks, and mentoring L1 teams to meet SLA/OLA targets.
Top Skills: Billing EngineCdr IngestionEsbJavaJavaScriptJIRALinuxMediation PlatformsNetcrackerNetcracker Convergent Billing & RatingProduct CatalogProduct Offering (Po)Product Structure (Ps)Rating EngineRemedyRest ApisRule EngineServicenowShell ScriptingSoap ApisSQLSsl CertificatesUnix

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account