You belong to the top echelon of talent in your field. At one of the world's most iconic financial institutions, where infrastructure is of paramount importance, you can play a pivotal role in shaping the reliability and resilience of systems that matter.
As a Lead Site Reliability Engineer at JPMorganChase within the Corporate Sector – Infrastructure Platforms, you hold a leadership role on your team, demonstrating strong knowledge across multiple technical domains and advising others on the technical and business challenges they face. You will lead resiliency design reviews, break complex problems into digestible work for other engineers, act as a technical lead for medium to large-scale products, and provide mentorship that elevates the entire team.
Job responsibilities
- Model and champion site reliability culture and practices; document and share knowledge across your organization through internal forums and communities of practice
- Lead initiatives to improve the reliability and stability of applications and platforms using data-driven analytics to improve service levels, proactively identifying and resolving technology-related bottlenecks
- Drive collaboration with your team to identify comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgets
- Serve as the primary point of contact during major incidents, applying deep technical expertise to identify and resolve issues quickly and minimize business impact
- Apply enterprise-authorized AI capabilities to accelerate major-incident triage, troubleshooting, post-incident analysis, and capacity risk identification — validating outputs and handling operational data according to sensitivity and security requirements
- Demonstrate an AI-first mindset by building and championing agentic automation for operational workflows (e.g., triage, runbook execution, incident summarization, change validation) with appropriate controls and monitoring
- Lead reuse-first adoption of AI-assisted reliability workflows across the software development lifecycle and toolchain practices, ensuring traceability, auditability, resiliency, and security controls
- Collect and analyze monitoring and telemetry data across test and production environments; design and implement dashboards to ensure system reliability, performance, and security
- Escalate issues with detailed technical write-ups; partner with application and infrastructure teams to identify and remediate capacity risks and understand platform interdependencies
- Offer a high level of technical expertise within one or more technical domains and provide advice and mentorship to other engineers
Required qualifications, capabilities, and skills
- Formal training or certification on site reliability engineering concepts and 5+ years applied experience
- Demonstrated proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and site reliability best practices, with hands-on experience implementing these within a platform
- Fluency in at least one programming language (e.g., Python, Java/Spring Boot, .NET, Go, Shell Scripting), including hands-on experience applying AI-assisted automation and agentic patterns to engineering and operational workflows
- Proficiency in scripting and automation and infrastructure-as-code tools (e.g., Python, PowerShell, Ansible, Terraform) and experience with cloud technologies across public and private environments
- Demonstrated experience using enterprise-authorized AI capabilities to improve site reliability engineering workflows (e.g., incident investigation, knowledge capture, capacity analysis) with strong validation habits and awareness of data sensitivity
- Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations
- Advanced knowledge and experience in observability, monitoring, alerting, and telemetry collection using tools such as Grafana, Dynatrace, Datadog, Prometheus, CloudWatch, or Splunk — including designing and implementing effective production monitoring dashboards
- Proficiency with continuous integration and continuous delivery practices and tooling, as well as container and container orchestration technologies
- Experience troubleshooting common networking technologies and issues, with knowledge of infrastructure areas including operating systems (Linux/Windows), databases, and deployment practices
- Advanced knowledge of software applications and technical processes with emerging depth in one or more technical disciplines, and a demonstrated drive to self-educate and evaluate new technologies
Preferred qualifications, capabilities, and skills
- Hands-on experience and certifications in AWS, Azure, GCP, or other cloud environments, with understanding of resiliency, scalability, observability, and monitoring
- Experience implementing CI/CD pipelines, conducting code reviews using GitHub, and building process automation with Python and scripting
- Experience utilizing Terraform or other infrastructure-as-code technologies for cloud resource management
- Proven ability to leverage GitHub Copilot or similar coding assistants to accelerate skill development and build AI agents or workflows that support infrastructure operations tasks (e.g., triage, runbook execution)
- Experience supporting complex, mission-critical applications involving multiple components across varying technical generations, with familiarity with modern front-end technologies
We offer a competitive total rewards package including base salary determined based on the role, experience, skill set and location. Those in eligible roles may receive commission-based pay and/or discretionary incentive compensation, paid in the form of cash and/or forfeitable equity, awarded in recognition of individual achievements and contributions. We also offer a range of benefits and programs to meet employee needs, based on eligibility. These benefits include comprehensive health care coverage, on-site health and wellness centers, a retirement savings plan, backup childcare, tuition reimbursement, mental health support, financial coaching and more. Additional details about total compensation and benefits will be provided during the hiring process.
We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.
JPMorgan Chase & Co. is an Equal Opportunity Employer, including Disability/Veterans
Similar Jobs
What you need to know about the Colorado Tech Scene
Key Facts About Colorado Tech
- Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
- Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
- Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
- Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
- Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

.png)
