Source Direct Logo

Source Direct

Principal Core Engineer — Infra / SRE

Reposted Yesterday
Be an Early Applicant
Hybrid
Denver, CO, USA
190K-215K Annually
Expert/Leader
Hybrid
Denver, CO, USA
190K-215K Annually
Expert/Leader
Own fleet-scale infrastructure and SRE architecture for an edge platform supporting thousands of devices. Lead reliability, scalability, secure lifecycle management, observability, incident response, upgradeability, and operational standards. Serve as technical owner during high-severity incidents, drive root-cause remediation, apply AI to operations, and coordinate cross-domain initiatives spanning software, hardware, networking, security, data, and AI runtime. Mentor senior engineers and provide architectural leadership across the organization.
The summary above was generated by AI
Principal Infrastructure / SRE Engineer

Hybrid — Denver, CO
Full-time

The Opportunity

We’re looking for a Principal Infrastructure / SRE Engineer to own the reliability, scalability, upgradeability, and operational excellence of an edge platform operating at fleet scale.

In this role, you’ll be the technical authority for designing and operating compound capabilities that span software, infrastructure, networking, security, data, and hardware. You’ll help ensure that fleets of thousands of devices can be reliably deployed, upgraded, and managed with the highest level of technical rigor.

You’ll set and enforce production standards and have the authority to stop changes that could put fleet safety or reliability at risk. During high-severity incidents, you’ll serve as the technical owner, leading root-cause analysis and driving durable fixes across teams.

This is a hands-on role for someone who thrives in a high-ownership environment and wants to build the infrastructure that makes real-world AI possible. You’ll work in an AI-native way, using AI to assist with diagnostics and operations while ensuring all production changes remain governed, reviewed, and auditable.

What You’ll Do
  • Own platform-wide reliability and scalability architecture across the fleet, including upgradeability, rollback safety, resilience, observability, and incident response.

  • Lead the design and delivery of compound capabilities spanning multiple specialist domains, including hardware, networking, security, data, infrastructure, and AI runtime.

  • Set and enforce production-grade standards for operational excellence, including SLOs/SLIs, error budgets, on-call readiness, change management, incident management, and postmortem practices.

  • Maintain the authority to stop changes that introduce unacceptable operational or fleet-level risk.

  • Serve as the technical owner during high-severity incidents, leading diagnosis, root-cause analysis, and coordinated remediation across teams.

  • Design and operate secure, automated fleet lifecycle systems for deployment, updates, configuration management, and health management at scale.

  • Drive the evolution of observability and telemetry systems, including metrics, logs, traces, audit data, and fleet state, so issues are detectable, diagnosable, and preventable.

  • Partner with engineering and commercial teams to translate real-world constraints into platform-level requirements and prioritization decisions.

  • Develop and use AI systems to accelerate diagnostics, automate operational workflows, and increase engineering velocity while ensuring production pathways remain governed, reviewed, and auditable.

  • Mentor senior engineers across domains, review technical designs, and raise the quality bar for architecture and reliability across the organization.

What Success Looks LikeIn your first 3 months, you will have:
  • Taken full ownership of a platform-wide reliability, upgradeability, or incident-reduction initiative and delivered measurable improvements in fleet stability, deployment safety, and operational clarity.

  • Established or strengthened production standards that reduce risk and improve consistency across releases and fleet operations.

  • Demonstrated strong incident ownership by leading at least one high-severity investigation through root cause and durable remediation.

In your first year, you will be:
  • Owning the fleet-scale operational architecture end-to-end, with clear accountability for reliability, upgradeability, scalability, and security posture across thousands of deployed systems.

  • Delivering significant improvements in platform resilience and operational excellence through durable systems, including automated lifecycle management, observability, incident reduction, and reliability standards.

  • Raising engineering rigor across the organization by enforcing standards, mentoring technical leaders, and driving cross-domain architectural decisions that compound over time.

Who You Are
  • 10+ years building and operating production infrastructure and distributed systems, including reliability engineering at scale across complex, multi-tenant, or fleet environments.

  • Deep experience with SRE practices, including SLOs/SLIs, error budgets, observability, incident response, postmortems, and operational automation.

  • Experience with Kubernetes-based platforms, Linux systems, and infrastructure-as-code automation.

  • Strong systems thinking across software, infrastructure, networking, and security, with the ability to drive outcomes across multiple domains and enforce production standards.

  • Proven ability to lead ambiguous, high-impact initiatives end-to-end with strong technical judgment, crisp execution, and disciplined change management.

  • Clear communicator and trusted technical partner to engineering leadership, with the ability to lead high-severity incident response and drive cross-team alignment.

  • Ownership mindset focused on outcomes rather than tasks.

Unique Experiences We Value
  • Designing and operating fleet management and upgrade systems at scale, including safe rollout and rollback, configuration management, and health monitoring.

  • Experience with canary deployments, staged rollouts, and verifiable rollback mechanisms.

  • Building observability platforms that make complex systems diagnosable and measurable across large distributed deployments.

  • Experience with metrics, logs, tracing pipelines, alerting, and dashboards that drive operational action.

  • Security-first operations experience involving secure boot, signed updates, audit logging, default-deny postures, and governed production changes.

  • Experience operating systems under real-world edge constraints, including limited connectivity, bandwidth limitations, variable environments, and high reliability requirements.

  • Building automation that reduces operational variance across large fleets.

  • Applying AI to operations and engineering workflows, including automated diagnostics, agentic triage, runbook generation, and anomaly detection, while keeping production pathways reviewed and auditable.

Benefits & Compensation
  • Work in a high-ownership, real-world startup environment where you can move quickly, build new systems, and see your impact directly through systems operating in the field.

  • Use modern AI tools throughout development, documentation, planning, troubleshooting, and operational workflows to accelerate execution.

  • Take on challenging technical problems across next-generation cloud and IoT, hardware/software/networking in real-world edge environments, data and AI inference, and secure systems operating in demanding OT settings.

  • Learn quickly by working with exceptional teammates and collaborating directly with industry leaders across software, AI, and infrastructure.

  • Base salary range of $190,000–$215,000, depending on location, experience, and comparable internal compensation.

  • Eligibility for meaningful equity through stock options in an early-stage, high-growth company.

  • Eligibility to participate in company benefit plans, which may include health, dental, and vision coverage, a 401(k) with company match, flexible PTO, paid parental leave, commuter benefits, and relocation and visa support for eligible roles.

Similar Jobs

6 Days Ago
Hybrid
Denver, CO, USA
190K-215K Annually
Expert/Leader
190K-215K Annually
Expert/Leader
Artificial Intelligence • Information Technology • Software • Infrastructure as a Service (IaaS)
Own fleet-scale reliability, upgradeability, and operational excellence for an edge platform. Design and operate automated, secure lifecycle systems, observability, and incident response. Lead cross-domain, high-severity incident ownership, set production standards (SLOs/SLIs, change management), mentor engineers, and apply AI to accelerate diagnostics and operational workflows.
Top Skills: Ai SystemsAudit LoggingCanary DeploymentsConfiguration ManagementFleet ManagementInfrastructure-As-CodeKubernetesLinuxObservabilitySecure BootStaged Rollouts
Yesterday
Hybrid
Littleton, CO, USA
16-24 Hourly
Junior
16-24 Hourly
Junior
eCommerce • Healthtech • Pet • Retail • Pharmaceutical
Provides front-desk and client-facing support in a veterinary practice. Responsibilities include patient intake, pet parent communication, customer service, issue resolution, PIMS use, following clinical and operating procedures, pharmacy responsibility, inventory communication, workspace cleaning, and required training. The role supports veterinarians, technicians, assistants, and practice management while delivering empathetic, high-quality care experiences.
Top Skills: Practice Information Management System (Pims)
Yesterday
Hybrid
Lakewood, CO, USA
15-24 Hourly
Junior
15-24 Hourly
Junior
eCommerce • Fashion • Retail • Sales • Wearables • Design
Serves as a Coach brand ambassador by delivering personalized luxury retail service, building client relationships, achieving individual and store sales goals, and using cross-selling, clienteling, mobile POS, and social selling. Supports transactions, inventory processing, replenishment, visual merchandising, online pickups, stockroom organization, and asset protection while collaborating with teammates. Requires flexible availability, strong communication, and the ability to lift and maneuver merchandise.
Top Skills: IpadLaptopMobile PosPoint-Of-Sale (Pos) Systems

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account