GE Vernova Logo

GE Vernova

System Reliability Engineering Lead

Posted An Hour Ago
Be an Early Applicant
Remote
Hiring Remotely in USA
152K-228K Annually
Expert/Leader
Remote
Hiring Remotely in USA
152K-228K Annually
Expert/Leader
Leads system reliability engineering for a global grid software SaaS portfolio. Owns cloud infrastructure, platform standardization, SLOs, release governance, incident command, disaster recovery, FinOps, and capacity planning. Serves as final production deployment authority and customer-facing reliability lead. Builds a reliability enablement center, mentors a distributed SRE team, and drives automation, progressive delivery, observability, compliance, and operational excellence across critical utility applications.
The summary above was generated by AI
Job Description SummaryAs the Tech Lead for System Reliability Engineering within the GridOS SaaS Products organization, you will be the hands-on technical authority on production stability for our global grid software SaaS portfolio. You will bridge the gap between architectural design and real-world operations, driving a culture of high reliability and engineering excellence across a distributed team spanning three geographies. You are the "Gatekeeper" for production environments — owning the Change Management process, holding final authority to approve or halt deployments based on system health, and accountable for meeting SLA/SLO targets for critical infrastructure applications serving major North American utility customers.
This is a player-coach role. You will architect and build alongside your team while setting technical direction, mentoring engineers, and serving as the primary customer-facing SRE point of contact. You will own FinOps for the SaaS platform, driving cloud cost optimization and capacity planning as the customer base scales.

Job DescriptionDay 0 — Strategic Provisioning and Design

Standardized Cloud Infra Provisioning

Architect and implement standardized, secure cloud infrastructure provisioning. Drive extreme automation to reduce account provisioning timelines and accelerate customer onboarding to the SaaS platform.

The Golden Path

Define and build the standardized "Middle-Mile" software delivery platform (IDP) using Backstage, ArgoCD, and GitHub Actions. Eliminate bespoke deployment methodologies and establish a single, repeatable path to production.

Follow-the-Sun Architecture

Design and operate the global handover protocols and 24/7 operational coverage model across US, India, and Mexico time zones. Ensure seamless support continuity without graveyard shifts.

Reliability Targets

Establish and own enterprise-wide Service Level Objectives (SLOs) and Service Level Indicators (SLIs) aligned with critical user journeys for global utility customers. Define error budgets and enforce them.

Day 1 — Release Governance and Deployment

Final Approval Authority

Serve as the final technical authority for all production releases. Enforce rigorous change control and validate that all security and performance quality gates are met before any deployment proceeds.

Progressive Delivery

Implement and operate advanced deployment strategies including Canary and Blue/Green rollouts. Build and verify automated rollback capabilities. Hands-on with deployment tooling and pipeline configuration.

SRE Center for Enablement (C4E)

Build and mature the C4E to provide coaching, standardized templates, and repeatable reliability patterns that uplift practices across all product teams. Act as the go-to technical resource for reliability engineering across the organization.

Day 2 — Operational Excellence and Optimization

Incident Command

Serve as the Lead Incident Commander for high-severity (Sev1/Sev2) events. Lead the technical direction, communication, and containment efforts. Available for P1 escalations around the clock.

Blameless Culture

Own the post-incident lifecycle. Facilitate blameless Root Cause Analysis (RCA) to ensure systemic fixes replace recurring operational risks. Build a team culture where incidents drive improvement, not blame.

Business Continuity

Architect and validate end-to-end Backup and Disaster Recovery (DR) strategies, including cross-region failover and automated recovery testing. Hands-on with DR runbook development and execution.

FinOps and Capacity Planning

Own financial operations for the SaaS platform. Drive cloud cost optimization through reserved instances, right-sizing, and waste elimination. Perform long-term capacity planning based on customer growth trajectory and application scaling requirements.

Customer Engagement and Team Leadership

Customer-Facing Accountability

Serve as the primary SRE point of contact for North American utility customers. Own customer satisfaction and NPS for SaaS reliability. Participate in customer-facing reviews, incident communications, and service health reporting. Must meet customer-mandated background check requirements for access to critical infrastructure data and environments.

Player-Coach Team Leadership

Lead a distributed team of 8 SRE engineers across Hyderabad Technical Center and Querétaro, scaling with SaaS application and customer growth. Set technical direction, assign tasks, own team deliverables, and drive day-to-day execution. Mentor engineers on SRE practices, cloud architecture, and operational discipline. Provide performance feedback to the people leader of record. Foster a culture of high performance and continuous learning.



Required Qualifications

Technical Qualifications

  • Cloud Ecosystem: Deep expertise in AWS core services (EC2, EKS, RDS, S3, IAM) and management tools (CloudTrail, CloudWatch)
  • Orchestration: Advanced mastery of Kubernetes internals and EKS cluster operations across multi-region architectures
  • Continuous Delivery: Expert knowledge of ArgoCD, GitHub Actions, and GitOps-first workflows
  • Automation: Proficiency in Infrastructure as Code (IaC) using Terraform and configuration management via Ansible
  • Observability: Hands-on experience with Prometheus, Grafana, observability platforms (Splunk or Datadog), and OpenTelemetry standard to build comprehensive telemetry pipelines
  • FinOps: Demonstrated experience in cloud cost optimization, reserved instance management, right-sizing, and long-term capacity planning for multi-tenant SaaS platforms

Experience and Leadership

  • Overall Experience: 12+ years in software engineering, cloud operations, or infrastructure roles
  • Domain Depth: 8–10 years of hands-on experience in SRE, Platform Engineering, Cloud Operations, or Production Support for large-scale, distributed SaaS applications
  • Technical Leadership: Proven track record of leading distributed engineering teams as a player-coach — setting technical direction while remaining hands-on with architecture, automation, and incident response
  • Operational Discipline: Exceptional troubleshooting skills under pressure and a "Fire Marshal" mindset toward investigation and proactive inspection
  • Customer Engagement: Experience working directly with enterprise customers on production reliability, incident communication, and service-level reporting
  • Background Check: Must be able to pass customer-mandated background screening for access to critical infrastructure environments


Desired Characteristics

Regulated Environments

  • Practical knowledge of NERC CIP compliance standards in a SaaS context
  • Experience with SOC2, ISO 27001, or IEC 62443 compliance frameworks
  • Familiarity with operating in highly regulated industries such as utilities, financial services, or critical national infrastructure

Certifications

  • AWS Certification: DevOps Engineer — Professional or Solutions Architect — Associate/Professional
  • CKA: Certified Kubernetes Administrator
  • SRE Practitioner Certification
  • AWS FinOps Practitioner or equivalent cloud financial management certification

Additional Information

About the SRE Team: The Grid Software SRE function is a newly established capability supporting the organization's SaaS transformation. The team currently supports Field Damage Assessment (FDA) and Distributed Dynamic Line Rating (DDLR) applications, with the portfolio expanding as GE Vernova Grid Software scales from its initial SaaS customers to a target of 20+ customers by end of 2027. The team operates a follow-the-sun model with engineers based in Hyderabad Technical Center (India) and Querétaro (Mexico).

Why US-Based: North American utility customers operating critical national infrastructure require that production environments and customer data be managed by US-based personnel who have completed customer-mandated background screening. This role exists to meet that requirement while providing hands-on technical leadership to the global SRE team.


Work Schedule

General shift, US business hours. On-call availability required for P1/Sev1 incidents. Follow-the-sun handoff protocols with Hyderabad and Querétaro teams.

Travel Requirements

Up to 10% — customer sites and team locations as needed (estimated 2–4 trips per year)


Additional Information

GE Vernova offers a great work environment, professional development, challenging careers, and competitive compensation. GE Vernova is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, national or ethnic origin, sex, sexual orientation, gender identity or expression, age, disability, protected veteran status or other characteristics protected by law.

GE Vernova will only employ those who are legally authorized to work in the United States for this opening. Any offer of employment is conditioned upon the successful completion of a drug screen (as applicable).

Relocation Assistance Provided: No

#LI-Remote - This is a remote position

For candidates applying to a U.S. based position, the pay range for this position is between $151,800.00 and $227,700.00. The Company pays a geographic differential of 110%, 120% or 130% of salary in certain areas. The specific pay offered may be influenced by a variety of factors, including the candidate’s experience, education, and skill set.

Bonus eligibility: discretionary annual bonus.

This posting is expected to remain open for at least seven days after it was posted on September 25, 2026.

Available benefits include medical, dental, vision, and prescription drug coverage; access to Health Coach from GE Vernova, a 24/7 nurse-based resource; and access to the Employee Assistance Program, providing 24/7 confidential assessment, counseling and referral services. Retirement benefits include the GE Vernova Retirement Savings Plan, a tax-advantaged 401(k) savings opportunity with company matching contributions and company retirement contributions, as well as access to Fidelity resources and financial planning consultants. Other benefits include tuition assistance, adoption assistance, paid parental leave, disability benefits, life insurance, 12 paid holidays, and permissive time off.

GE Vernova Inc. or its affiliates (collectively or individually, “GE Vernova”) sponsor certain employee benefit plans or programs GE Vernova reserves the right to terminate, amend, suspend, replace, or modify its benefit plans and programs at any time and for any reason, in its sole discretion. No individual has a vested right to any benefit under a GE Vernova welfare benefit plan or program. This document does not create a contract of employment with any individual.

GE Vernova Boulder, Colorado, USA Office

208 Wild Tiger Rd, Boulder, Colorado, 80302-9263, United States, Boulder, United States, 80302-9263

Similar Jobs

A Minute Ago
Remote
Pennsylvania, USA
Mid level
Mid level
Healthtech • Logistics • Pharmaceutical
Coordinates information security human risk programs, including training deployments, simulated phishing campaigns, onboarding, targeted learning, awareness communications, metrics, and audit-ready documentation. Drafts behavior-focused educational materials, tracks initiative milestones and risks, reviews campaign results, identifies process improvements, and collaborates with Information Security, HR, Legal, Communications, and business stakeholders.
Top Skills: Learning Management SystemsMarketing CloudExcelMS OfficeMicrosoft PowerpointMicrosoft SharepointMicrosoft TeamsSimulated Phishing Platforms
A Minute Ago
Remote
Pennsylvania, USA
38K-57K Annually
Internship
38K-57K Annually
Internship
Healthtech • Logistics • Pharmaceutical
Supports Cencora’s enterprise data and analytics platform modernization by collecting and analyzing datasets, creating dashboards with BI tools, tracking KPIs, conducting ad hoc analyses, documenting findings, and contributing to CI/CD and self-service solution delivery. Collaborates with cross-functional stakeholders to provide actionable insights, improve processes, and support data-driven decisions for internal and customer-facing products.
Top Skills: Ci/CdExcelPower BISalesforceSQLTableau
A Minute Ago
Remote
Pennsylvania, USA
38K-57K Annually
Internship
38K-57K Annually
Internship
Healthtech • Logistics • Pharmaceutical
Supports warehouse automation projects involving material-handling controls, software and systems implementation, specifications, testing, issue tracking, documentation, deployment, training, and process improvement. Collaborates with project teams and business partners, analyzes trends, maintains project updates, and presents internship outcomes. The role is remote with likely distribution-center visits and is intended for students in supply chain, logistics, or engineering programs.
Top Skills: DatabasesDatabricksJavaMaterial Handling ControlsAzureExcelMS OfficeMicrosoft OutlookMicrosoft PowerpointMicrosoft TeamsPower BIPythonVisual Studio

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account