Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Denver & Boulder, CO
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead long-term strategy and architecture for cloud and on‑prem platform infrastructure, driving Kubernetes and multi‑cloud reliability, IaC/GitOps automation, observability, SLO/SLI/error‑budget practices, incident leadership, AI‑augmented tooling adoption, and mentorship of senior engineers to improve platform resilience and developer experience.
Top Skills:
Amazon Elastic Kubernetes Service (Eks)AutoscalingAWSCapacity PlanningCi/CdGitopsGoGoogle Cloud PlatformGoogle Kubernetes Engine (Gke)Identity And Access ManagementInfrastructure As CodeKubernetesLinuxNetworkingObservabilityOperatorsPulumiPythonRke2StorageTerraform
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills:
AnsibleAws EcsKubernetesLinuxPythonTerraform
Reposted 5 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
As a Cloud Cost Utilization SRE at GitLab, you'll manage cloud spending, improve tracking and optimization of cloud usage, and collaborate with finance and engineering teams to enhance cost efficiency across AWS and GCP.
Top Skills:
AnsibleAWSElkGCPGrafanaLokiMimirPrometheusTempoTerraform
Digital Media • Information Technology • News + Entertainment
Senior SRE responsible for capacity planning, incident response, monitoring and alerts, performance tuning, disaster recovery, cloud infrastructure optimization, automation, and cross-functional collaboration to ensure scalable, reliable services.
Top Skills:
AnsibleAWSAzureBashDatadogDockerGCPKubernetesNoSQLPythonSplunkSQLTerraform
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Artificial Intelligence • Machine Learning
Lead development of AI-assisted reliability tooling, own incident response end-to-end, improve observability and SLO/SLI frameworks, scale single-tenant SaaS operations, mentor engineers, and reduce recurring operational toil through engineering and automation.
Top Skills:
Cloud PlatformsGoKubernetesLinuxLlm/Ai ToolingLogs And TracingObservability ToolingPythonSlo/Sli Frameworks
Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
The Staff Site Reliability Engineer will lead reliability strategies, manage high-risk initiatives, and enhance engineering standards while ensuring system reliability and operational excellence within a hybrid work environment.
Top Skills:
BashCi/CdDatabase ArchitectureGoGoogle Cloud PlatformInfrastructure-As-CodeKubernetesMonitoring PlatformsPulumiPythonTerraform
12 Days AgoSaved
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills:
AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Reposted 12 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Cloud • Information Technology • Security • Software • Cybersecurity
As a Staff Site Reliability Engineer, you'll oversee Zscaler production data center services, optimize code, and ensure cloud service availability and performance. Collaborate with cross-functional teams to improve processes and resolve escalated issues.
Top Skills:
BashDnsFirewallsGrafanaHTTPIcmpLoad BalancingNagiosOsi ModelPrometheusPythonTcp/Ip
Reposted 15 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills:
.NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Build and maintain automation and reliability for live video distribution across on-prem and cloud. Deploy and manage systems, develop monitoring and automated recovery, troubleshoot complex incidents, coordinate with vendors, document SOPs, support live broadcast components, and participate in L2 on-call rotation.
Top Skills:
AacAc3AnsibleAtscAvcAWSBashChefCloudFormationCmafDockerEksGitHevcHlsJavaScriptJSONKubernetesLinuxMicrosoft Graph ApiMpeg Transport StreamsPythonRistScte104Scte224Scte35SrtSsaiSt2022-7St2110StatmuxTerraformUnixXMLYmlZixi
Reposted 21 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills:
AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills:
AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
Information Technology • Insurance • Software
Own and operate production services end-to-end to ensure reliability, scalability, performance, and operational health. Define SLIs/SLOs, perform incident response and root cause analysis, build automation and self-healing, manage production changes, and collaborate with engineering, product, and operations teams to improve system design and observability.
Top Skills:
.NetAWSC#Ci/CdInfrastructure As CodeJavaKubernetesLinuxPythonReactRelational DatabasesWindows
Information Technology • Insurance • Software
Define and own enterprise reliability, scalability, and performance for production services. Drive architectural standards, observability strategy, SLO/SLI and error-budget governance, lead incident command for high-severity events, and foster a blameless, engineering-first operations culture across cloud, hybrid data centers, and customer-hosted environments.
Top Skills:
.NetAWSC#Ci/CdInfrastructure-As-CodeJavaKubernetesLinuxObservabilityPythonReactRelational DatabasesWindows
Healthtech • Telehealth
Lead the creation of SRE practices across six product teams: define SLOs/SLIs, implement observability, reduce outages, automate toil with IaC and tooling, run production readiness reviews, and train teams to improve incident detection and reliability.
Top Skills:
AWSAws EksDatadogGrafanaKubernetesNode.jsPrometheusPythonRdsReactTerraformTypescript
Software
Join as the company's first SRE to design reliability processes, build observability/CI/CD/IaC tooling, define SLOs, embed with product teams, run incident response, and support compliance for patient data.
Top Skills:
Access ControlsAlerting)Audit TrailsCi/CdGoHipaaInfrastructure As CodeLoggingObservability (MetricsPythonSlosSoc 2TracingTypescript
Other • Social Impact
The Senior Site Reliability Engineer is responsible for maintaining Wikimedia's infrastructure, improving reliability, automating processes, and collaborating with teams. The role involves troubleshooting, managing deployments, and leading incident responses while working remotely.
Top Skills:
AnsibleBashCassandraDebianGoGrafanaHhvmKubernetesMariadbMemcachedPHPPrometheusPuppetPythonRedisRubyShell
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead qualification, performance validation, and production readiness for new Azure Storage hardware and firmware. Drive test planning, automation frameworks, large-scale telemetry and benchmark analysis, root-cause investigations across software/hardware/firmware, and partner with engineering and vendors to resolve reliability and performance issues.
Top Skills:
Automation FrameworksAzureAzure StorageBenchmarkingFirmwareNetworkingSsdsTelemetry
Legal Tech • Software
Lead observability and incident management efforts: define SLIs/SLOs, build monitoring/alerting, dashboards, logging, and tracing. Drive incident response, postmortems, and reliability improvements to reduce MTTD/MTTR. Integrate observability into CI/CD, maintain AWS and Kubernetes infrastructure, automate operations, and mentor engineers on SRE best practices.
Top Skills:
AWSBashCi/CdDatadogDistributed TracingDynatraceGrafanaKubernetesNew RelicOpentelemetryPowershellPrometheusPython
Reposted YesterdaySaved
Legal Tech • Software
Own and improve platform reliability, availability, and performance for Filevines systems. Build AWS infrastructure with Terraform, develop CI/CD pipelines, monitoring, and automation, define SLOs/SLIs/error budgets, lead incident response, mentor engineers, and participate in a 24/7 on-call rotation to support highly available, low-latency production systems.
Top Skills:
AutomationAWSBashCi/CdCloudwatchEc2EcsEksIamLambdaObservabilityPowershellPythonRoute 53S3TerraformVpc
Legal Tech • Software
Design and improve observability (monitoring, logging, tracing, SLIs/SLOs), build automation and CI/CD, lead incident response and reliability improvements, mentor SREs, run on-call, and apply AI/ML to operational signals to forecast and reduce risks.
Top Skills:
AWSBashCi/CdDistributed TracingGoInfrastructure As CodeKubernetesLoggingMonitoringPythonSlisSlos
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Denver & Boulder, CO Companies Hiring SRE Engineers
See AllPopular Denver & Boulder, CO Engineering Job Searches
Engineering Jobs in Denver & Boulder, CO
Software Engineer Jobs in Denver & Boulder, CO
Android Developer Jobs in Denver & Boulder, CO
C# Jobs in Denver & Boulder, CO
C++ Jobs in Denver & Boulder, CO
DevOps Jobs in Denver & Boulder, CO
Front End Developer Jobs in Denver & Boulder, CO
Golang Jobs in Denver & Boulder, CO
Hardware Engineer Jobs in Denver & Boulder, CO
iOS Developer Jobs in Denver & Boulder, CO
Java Developer Jobs in Denver & Boulder, CO
Javascript Jobs in Denver & Boulder, CO
Linux Jobs in Denver & Boulder, CO
Engineering Manager Jobs in Denver & Boulder, CO
.NET Developer Jobs in Denver & Boulder, CO
PHP Developer Jobs in Denver & Boulder, CO
Python Jobs in Denver & Boulder, CO
QA Jobs in Denver & Boulder, CO
Ruby Jobs in Denver & Boulder, CO
Salesforce Developer Jobs in Denver & Boulder, CO
Scala Jobs in Denver & Boulder, CO
Associate Software Engineer Jobs in Denver & Boulder, CO
Automation Engineer Jobs in Denver & Boulder, CO
Backend Engineer Jobs in Denver & Boulder, CO
Cloud Engineer Jobs in Denver & Boulder, CO
Controls Engineer Jobs in Denver & Boulder, CO
CTO Jobs in Denver & Boulder, CO
Design Engineer Jobs in Denver & Boulder, CO
DevOps Engineer Jobs in Denver & Boulder, CO
Director of Engineering Jobs in Denver & Boulder, CO
Electrical Engineering Jobs in Denver & Boulder, CO
Embedded Software Engineer Jobs in Denver & Boulder, CO
Full-Stack Engineer Jobs in Denver & Boulder, CO
Infrastructure Engineer Jobs in Denver & Boulder, CO
Manufacturing Engineer Jobs in Denver & Boulder, CO
Mechanical Design Engineer Jobs in Denver & Boulder, CO
Mechanical Engineering Jobs in Denver & Boulder, CO
Network Engineer Jobs in Denver & Boulder, CO
Platform Engineer Jobs in Denver & Boulder, CO
Principal Engineer Jobs in Denver & Boulder, CO
Principal Software Engineer Jobs in Denver & Boulder, CO
Process Engineer Jobs in Denver & Boulder, CO
Project Engineer Jobs in Denver & Boulder, CO
QA Engineer Jobs in Denver & Boulder, CO
Robotics Engineer Jobs in Denver & Boulder, CO
Security Engineer Jobs in Denver & Boulder, CO
Software Architect Jobs in Denver & Boulder, CO
Solutions Architect Jobs in Denver & Boulder, CO
Solutions Engineer Jobs in Denver & Boulder, CO
SRE Engineer Jobs in Denver & Boulder, CO
Staff Software Engineer Jobs in Denver & Boulder, CO
Systems Engineer Jobs in Denver & Boulder, CO
All Filters
Total selected ()
No Results
No Results











.png)
.png)

















