Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Denver & Boulder, CO
Software • Consulting
Lead 24x7 application support for external web applications: manage incidents, perform RCA, implement preventative fixes, build monitoring/alerts, expand Splunk functionality, create dashboards, collaborate with development and platform teams, and participate in on-call rotation.
Top Skills:
ApmAppdynamicsAWSDatadogGrafanaKubernetesLinuxMulesoftOpenshiftOpentelemetryPostmanPythonRumSeleniumServicenowShell ScriptingSplunk (Spl)Splunk CloudSplunk Observability CloudSplunk Synthetics
Aerospace • Defense
Own and operate Loft's cloud and hybrid network infrastructure: design architectures, secure connectivity, manage network-as-code (IaC), define SLOs and observability, automate toil, and participate in SRE incident response and platform reliability.
Top Skills:
ArgocdCi/CdDnsDockerFirewallingFluxcdGCPGitGitopsGrafanaHybrid ConnectivityK8SKubernetesPeeringSdnSite-To-Site RoutingSoftware-Defined NetworkingTerraformVpcVpn
Information Technology • Cybersecurity • Defense • Automation
Design, build, and maintain secure, highly available Azure cloud-native platforms and CI/CD pipelines for classified mission systems. Implement Infrastructure as Code, automation, monitoring, and security controls while supporting hybrid Windows environments and platform reliability.
Top Skills:
Arm TemplatesAzureAzure ComputeAzure DevopsAzure IdentityAzure Kubernetes ServiceAzure MonitorAzure NetworkingAzure StorageBicepC#Ci/CdDevsecopsDockerGitInfrastructure As CodeKubernetesLog AnalyticsPowershellPythonTerraformWindows Server
Software
Lead SRE to define SRE strategy, architecture, and roadmap; design and operate containerized, compliant cloud environments; build observability, incident management, automation, and developer platform capabilities; mentor SRE team and collaborate with security, compliance, and product teams to ensure reliability at scale.
Top Skills:
AWSAws MarketplaceAzureAzure MarketplaceGCPGoogle Cloud MarketplaceGrafanaKubernetesPrometheusTerraform
Artificial Intelligence • Other • Sales • Software
The role involves designing and advancing infrastructure for the engineering team, ensuring the reliability of Kubernetes clusters, automating operations, and building machine learning infrastructure.
Top Skills:
ArgoAWSAzureCloudFormationFluxGithub ActionsGoGCPKubernetesPostgresPythonTerraform
Agency • Information Technology
Lead SRE role designing and maintaining CI/CD pipelines (GitHub Actions), containerized deployments (Docker, Kubernetes, AKS, Helm), web/mobile app releases, observability, automated testing, and DevOps best practices across cloud environments with cross-functional collaboration and regulatory compliance.
Top Skills:
AksAndroidAzure Application InsightsAzure Log AnalyticsAzure MonitorBashBranchingDockerDocker ComposeGitGit HooksGithub ActionsGoogle PlayHelmHerokuiOSIos App StoreJavaKubernetesNpmPowershellPull RequestsPythonSonarqubeVeracodeVercel
Artificial Intelligence • Insurance • Software • Automation
The Staff Site Reliability Engineer will build and scale infrastructure for Assured's platform, automate delivery, enhance observability, and lead mentoring initiatives.
Top Skills:
AWSKubernetesPostgresTerraform
Reposted 3 Days AgoSaved
Aerospace • Information Technology • Professional Services • Security • Software
Maintain and improve reliability, scalability, and performance of enterprise infrastructure across global sites. Implement automation and infrastructure-as-code, build monitoring and observability, perform RCA and incident response, support patching and RMF changes, integrate new capabilities, and maintain operational documentation and ITIL/ITSM processes to ensure mission-ready, high-availability environments.
Top Skills:
AnsibleElkNagiosPowershellPythonScomSolarwindsSplunkTerraform
Cloud • Security
Build and operate the production platform (Kubernetes, AWS, IaC, CI/CD, observability), automate self-service deployment, embed security and secrets management, run and modernize on-call, drive cost efficiency, mentor teammates, and maintain runbooks and post-incident reviews.
Top Skills:
AWSBashCi/CdClaudeGitGrafanaKubernetesLinuxPrometheusPythonSaltTerraform
Healthtech • Social Impact • Software
Own the operational lifecycle of cloud-native data infrastructure: design and automate reliable deployments, observability, incident response, SLIs/SLOs, autoscaling and IaC, and improve platform efficiency and data freshness across GKE and Cloud Run.
Top Skills:
BashBigQueryCloud BuildCloud MonitoringCloud RunDatadogDockerGCPGithub ActionsGkeGoGrafanaJIRAKubernetesPrometheusPulumiPythonSentrySlackSnykSonarqubeTerraform
Software
Own and improve platform performance, reliability, and deployment automation. Manage cloud infrastructure, implement IaC, monitor systems with observability tools, provide operational support for distributed applications, and integrate production learnings into development workflows.
Top Skills:
Aiops ToolingAws Elastic ContainersAws RdsAws S3Claude CodeClaude CoworkDatadogHarness EngineeringInfrastructure As CodeKubernetesLlmsPrompt EngineeringRigorSplunk
Cloud • Security • Software • Cybersecurity
Lead reliability for a serverless AI inference platform: own observability and SLO/SLI frameworks, build automation and tooling, manage incidents and on-call, define deployment safety (canaries, rollbacks), influence architecture with product teams, and mentor other SREs.
Top Skills:
AutoscalingCi/CdContainer OrchestrationContainerizationGoGpu WorkloadsInfrastructure-As-CodeKubernetesModel ServingPythonResource Scheduling
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Artificial Intelligence • Healthtech • Software • Telehealth
Design, deploy, and maintain AWS-hosted Kubernetes (EKS) infrastructure; build automation and AI-assisted runbooks; provide observability and incident response; ensure HIPAA-compliant, high-availability platform operations and mentor engineering teams.
Top Skills:
AWSBashDatadogEc2EksGitGithub ActionsGoHelmKubernetesPythonRdsS3Terraform
Other
The Senior Site Reliability Engineer at Juul Labs ensures operational stability and performance of hybrid cloud infrastructure, leads automation, and handles critical incidents.
Top Skills:
AWSBashCloudFormationGCPNutanixPowershellPythonTerraform
Digital Media • Social Media • Software • Sports
Lead the technical architecture and execution of migration to AWS, drive developer enablement, and automate infrastructure using code-first principles.
Top Skills:
Aws EksDatadogGithub ActionsGoIstioK6KubernetesNode.jsTerraform
Computer Vision • Machine Learning • Software
As a Site Reliability Engineer, ensure the reliability, performance, and scalability of Ditto's cloud infrastructure by developing observability solutions, leading incident management, and collaborating with product engineering teams.
Top Skills:
AWSAzureCDatadogGCPGoGrafanaHelmJavaKubernetesPrometheusRustTerraform
Artificial Intelligence • Fintech • Machine Learning • Natural Language Processing • Business Intelligence
Lead architecture and implementation of reliability platforms and SRE practices for a production SaaS. Build self-service reliability tooling, drive AIOps automation, advance observability (monitoring, tracing, profiling), lead incident response and postmortems, mentor engineers, and embed production readiness across teams to achieve 99.99% uptime.
Top Skills:
AWSAzureContinuous ProfilingDatadogDnsElkGCPGoGrafanaHttp/SKubernetesLoad BalancingOpentelemetryPrometheusPythonTcp/Ip
AdTech • Big Data • Marketing Tech • Software
Responsible for owning and optimizing the Internal Developer Platform, improving reliability, scalability, and usability while supporting engineering teams and standardizing operational processes through automation and best practices.
Top Skills:
ArmAWSAzureBashCloudFormationConsulDockerGithub ActionsHashicorpJenkinsKubernetesLinuxNomadPowershellPythonSplunkSumo LogicTerraformVaultWindows
Database • Analytics
This role involves ensuring the reliability and performance of ClickHouse's cloud infrastructure, collaborating with engineering teams, incident management, and driving continuous improvement in service availability.
Top Skills:
AnsibleAWSAzureClickhouseDocker SwarmGoGoogle Cloud PlatformKubernetesPuppetPythonTerraform
Software • Financial Services
Ensure platform reliability, performance, and availability by implementing observability, automating infrastructure, participating in on-call rotations and post-mortems, partnering with Product and Engineering, designing scalable architectures, mentoring teammates, and integrating Dynatrace with Azure DevOps and Jira while supporting compliance (SOC/FedRAMP).
Top Skills:
.NetAksAlpineAnsibleAppinsightsArm TemplatesAWSAzure DevopsBashBicepC#ChefCloudFormationDatadogDebianDynatraceEksGCPGitGitGksGrafanaHelmJIRAKubernetesLog AnalyticsAzureNew RelicOnestream SoftwareOpenshiftPowershellPowershell DscPrometheusPuppetPythonRest ApisSQLTerraformUbuntu
Fintech • Information Technology
As a Site Reliability Engineer at Alpaca, you will ensure system reliability and performance, troubleshoot issues, and collaborate with teams to design scalable features.
Top Skills:
GoGormLinuxPgxPostgresPrometheusSqlc
Artificial Intelligence • Cloud • Events • Productivity • Software • Business Intelligence • Conversational AI
Maintain and improve uptime, availability, and performance of services via observability, redundancy, failover, and load‑balancing. Integrate monitoring into SDLC, lead incident response/on‑call, assess capacity and risks, and work with teams to extend observability and automate self‑healing.
Top Skills:
AlertmanagerAnsibleArgocdAWSAzureBashElkGCPGitlabGitlab CiGoGrafanaJavaJavaScriptJenkinsKafkaKubernetesLinuxMongoDBMySQLNginxPostgresPrometheusPythonTerraformVictoriametricsZabbix
Healthtech
Design, scale, and operate secure AWS cloud infrastructure (EKS, IAM, RBAC); build and maintain IaC (Terraform/Terragrunt), GitHub Actions CI/CD, Datadog observability, and Python automation; document runbooks, participate in on-call rotations, postmortems, and Agile workflows to improve reliability and security.
Top Skills:
AWSDatadogEc2EksFargateGithub ActionsGithub Advanced SecurityHelmIamJIRAKubernetesLambdaPythonRbacSecrets ManagerServerlessTerraformTerragruntVpc
Gaming • Software
The Site Reliability Engineer will manage infrastructure stability and scalability, lead cloud migrations, and optimize performance across systems while mentoring team members.
Top Skills:
AnsibleAWSAzureBashChefCloudFormationDatadogDockerElk StackGCPGoGrafanaKubernetesPrometheusPuppetPythonTerraformUnix/Linux
Big Data • Analytics
Own production reliability for customer-facing radar and weather data services across Azure, colocation, and edge Kubernetes. Refactor C#/.NET services for multi-replica safety, design multi-cluster HA, operate self-managed Kubernetes, improve observability and automation, lead incident response and postmortems, and drive operational excellence and capacity planning.
Top Skills:
.NetAnsibleC#DatadogGpu-Enabled WorkloadsGrafanaHelmIstioKubernetesLokiLonghornAzureNatsOctopus DeployOpentelemetryPostgisPostgresPrometheusRabbitMQRancherRke2Terraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Denver & Boulder, CO Companies Hiring SRE Engineers
See AllPopular Denver & Boulder, CO Engineering Job Searches
Engineering Jobs in Denver & Boulder, CO
Software Engineer Jobs in Denver & Boulder, CO
Android Developer Jobs in Denver & Boulder, CO
C# Jobs in Denver & Boulder, CO
C++ Jobs in Denver & Boulder, CO
DevOps Jobs in Denver & Boulder, CO
Front End Developer Jobs in Denver & Boulder, CO
Golang Jobs in Denver & Boulder, CO
Hardware Engineer Jobs in Denver & Boulder, CO
iOS Developer Jobs in Denver & Boulder, CO
Java Developer Jobs in Denver & Boulder, CO
Javascript Jobs in Denver & Boulder, CO
Linux Jobs in Denver & Boulder, CO
Engineering Manager Jobs in Denver & Boulder, CO
.NET Developer Jobs in Denver & Boulder, CO
PHP Developer Jobs in Denver & Boulder, CO
Python Jobs in Denver & Boulder, CO
QA Jobs in Denver & Boulder, CO
Ruby Jobs in Denver & Boulder, CO
Salesforce Developer Jobs in Denver & Boulder, CO
Scala Jobs in Denver & Boulder, CO
Associate Software Engineer Jobs in Denver & Boulder, CO
Automation Engineer Jobs in Denver & Boulder, CO
Backend Engineer Jobs in Denver & Boulder, CO
Cloud Engineer Jobs in Denver & Boulder, CO
Controls Engineer Jobs in Denver & Boulder, CO
CTO Jobs in Denver & Boulder, CO
Design Engineer Jobs in Denver & Boulder, CO
DevOps Engineer Jobs in Denver & Boulder, CO
Director of Engineering Jobs in Denver & Boulder, CO
Electrical Engineering Jobs in Denver & Boulder, CO
Embedded Software Engineer Jobs in Denver & Boulder, CO
Full-Stack Engineer Jobs in Denver & Boulder, CO
Infrastructure Engineer Jobs in Denver & Boulder, CO
Manufacturing Engineer Jobs in Denver & Boulder, CO
Mechanical Design Engineer Jobs in Denver & Boulder, CO
Mechanical Engineering Jobs in Denver & Boulder, CO
Network Engineer Jobs in Denver & Boulder, CO
Platform Engineer Jobs in Denver & Boulder, CO
Principal Engineer Jobs in Denver & Boulder, CO
Principal Software Engineer Jobs in Denver & Boulder, CO
Process Engineer Jobs in Denver & Boulder, CO
Project Engineer Jobs in Denver & Boulder, CO
QA Engineer Jobs in Denver & Boulder, CO
Robotics Engineer Jobs in Denver & Boulder, CO
Security Engineer Jobs in Denver & Boulder, CO
Software Architect Jobs in Denver & Boulder, CO
Solutions Architect Jobs in Denver & Boulder, CO
Solutions Engineer Jobs in Denver & Boulder, CO
SRE Engineer Jobs in Denver & Boulder, CO
Staff Software Engineer Jobs in Denver & Boulder, CO
Systems Engineer Jobs in Denver & Boulder, CO
All Filters
Total selected ()
No Results
No Results





















%20(1).png)














