Blackpoint Cyber Logo

Blackpoint Cyber

Sr. Site Reliability Engineer

Posted Yesterday
Remote
Hiring Remotely in United States
150K-187K Annually
Senior level
Remote
Hiring Remotely in United States
150K-187K Annually
Senior level
Designs, operates, and improves scalable cloud and on-premise infrastructure, CI/CD pipelines, Kubernetes platforms, data streaming systems, and observability tooling. Owns AWS reliability, security, cost efficiency, automation, progressive deployments, incident response, and production troubleshooting. Partners with software engineering teams to integrate services and drive continuous improvements in system performance, scalability, and uptime.
The summary above was generated by AI

Blackpoint Cyber is the leading provider of world-class cybersecurity threat hunting, detection and remediation technology. Founded by former National Security Agency (NSA) cyber operations experts who applied their learnings to bring national security-grade technology solutions to commercial customers around the world, Blackpoint Cyber is in hyper-growth mode,  fueled by a recent $190m series C round. 

SUMMARY

We're hiring a Senior Site Reliability Engineer to design, implement, and maintain our cloud and on-premise infrastructure and CI/CD pipelines, with a focus on automation, scalability, and performance. You'll work across cloud platform administration, container orchestration, data streaming, observability, and incident response — partnering with engineering teams to keep our systems reliable, secure, and efficient, and helping foster a culture of continuous improvement.

RESPONSIBILITIES

  • Design, develop, and maintain highly scalable infrastructure using Infrastructure as Code (Terraform and Terragrunt) for automated cloud resource provisioning and orchestration.

  • Own and optimize our AWS cloud environment, ensuring cost efficiency, security best practices, and high-availability standards.

  • Manage and optimize Kubernetes cluster environments (Helm, ArgoCD, Istio, Kustomize) to support continuous delivery and infrastructure-as-code practices.

  • Administer and scale data streaming infrastructure (Confluent Cloud, Apache Kafka) to support enterprise-level data processing.

  • Deploy, configure, and maintain Redis for caching and real-time data processing.

  • Implement and maintain monitoring, alerting, and incident response frameworks (Prometheus, Grafana, Alert Manager, Grafana CloudOpsGenie/PagerDuty) to ensure system reliability and performance.

  • Facilitate controlled feature deployments and progressive rollouts through LaunchDarkly/PostHog.

  • Partner with software development teams to ensure seamless integration of new services, applications, and features into existing infrastructure.

  • Diagnose and resolve complex system-level issues, implementing solutions that maintain high performance and maximize uptime.

  • Drive continuous improvement of automation tooling, operational processes, and engineering methodologies to enhance scalability, reliability, and maintainability.

  • Stay current on emerging SRE trends and tools, andtools and help the team adopt relevant industry advancements and best practices.

REQUIREMENTS

  • 5+ years of experience in a Senior Site Reliability Engineer role or equivalent, with substantial emphasis on cloud infrastructure management and automation.

  • Expertise in Infrastructure as Code (Terraform, Terragrunt) for enterprise-scale deployments.

  • Comprehensive knowledge of AWS, including designing, implementing, and maintaining secure, scalable, resilient cloud architectures.

  • Extensive hands-on experience with distributed data streaming (Confluent Cloud, Apache Kafka).

  • Proven experience with Redis for caching and Amazon RDS for relational database management.

  • Experience with enterprise search and analytics platforms (OpenSearch, Elasticsearch, ChaosSearch).

  • Proficiency designing and implementing monitoring/alerting infrastructure (Prometheus, Grafana, Alert Manager, Grafana Cloud, OpsGenie/PagerDuty).

  • Practical experience with feature flag systems (LaunchDarkly/PostHog) for controlled release management.

  • Extensive experience administering production-grade Kubernetes (Helm, ArgoCD, Istio); working knowledge of Kustomize.

  • Strong problem-solving skills, with the ability to troubleshoot complex systems in production.

  • Strong communication and collaboration skills, with experience working in Agile environments.

NICE TO HAVE

  • Experience with Terragrunt to manage Terraform across multiple environments.

  • Extensive hands-on experience with distributed data streaming (Kafka).

  • Multi-cloud experience (Google Cloud Platform, Microsoft Azure).

  • Understanding of security frameworks and compliance standards for cloud-native/containerized environments.

  • Serverless computing and CI/CD pipeline experience (Jenkins, GitHub Actions).

  • Software development proficiency in Node.js, Python, and/or Go.

Blackpoint Cyber welcomes and encourages applications from qualified individuals of all races, colors, religions, sex, sexual orientation, gender identity or expression, national origin, age, marital status, or any other legally protected status. We are committed to equality of opportunity in all aspects of employment.

For eligible employees in the US, Blackpoint offers competitive Health, Vision, Dental, and Life Insurance plans, a robust 401k plan, Discretionary Time Off, and other minor perks. International employees receive competitive benefits in accordance with local market standards and applicable country requirements.

Blackpoint believes all employees should share in the company’s success – equity participation is available to employees globally, with program details varying by location and employment structure.

HQ

Blackpoint Cyber Denver, Colorado, USA Office

1099 18th St, Suite 3050, Denver, Colorado, United States, 80202

Similar Jobs

Yesterday
In-Office or Remote
75K-195K Annually
Senior level
75K-195K Annually
Senior level
Cloud • Software
Own NetBox’s build and release pipeline from image creation through Cloud and Enterprise deployment. Improve Django and PostgreSQL performance, establish observability and SLOs, strengthen software supply chain security, and support SOC 2 compliance. Participate in on-call rotations, incident response, and postmortems. Drive cross-team migrations and release processes while contributing fixes to NetBox Core when reliability issues originate in the application.
Top Skills: ArgocdAws Ec2Aws IamAws RdsAws VpcClaude CodeCosignDjangoFluxcdGithub ActionsGrafanaHelmKubernetesPostgresPrometheusPythonSigstoreSlsaTerraform
5 Days Ago
In-Office or Remote
92K-164K Annually
Senior level
92K-164K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Designs and operates secure, reliable Azure cloud platforms using Terraform, GitHub Actions, containers, and automation. Responsibilities include CI/CD, observability, incident response, platform security, vulnerability remediation, disaster recovery, infrastructure troubleshooting, and SRE practices. The role supports production workloads, improves reliability and delivery processes, participates in on-call activities, and mentors engineers while partnering across development, security, architecture, and operations teams.
Top Skills: BashCi/CdCloud SecurityDockerGitGithub ActionsGitopsInfrastructure As CodeKubernetesAzureObservabilityPowershellPythonTerraform
13 Days Ago
Remote or Hybrid
United States
Senior level
Senior level
Fintech • Software
The Senior Site Reliability Engineer ensures SaaS platforms remain reliable, performant, secure, and scalable. Responsibilities include building cloud infrastructure, implementing monitoring and alerting, automating operational runbooks and deployments, managing Infrastructure as Code, applying AI-powered observability and remediation, supporting Kubernetes and cloud networking, and leading incident triage and root-cause analysis during 24/7 on-call rotations.
Top Skills: AIAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC# .NetCi/CdCloud NetworkingCloudopsCosmos DbDatadogDynatraceEksFirewallsHarnessIdera Sql Diagnostic ManagerInfrastructure As CodeJavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account