Metasys Logo

Metasys

Site Reliability Engineer Internship

Reposted 11 Hours Ago
Remote
Hiring Remotely in United States
Internship
Remote
Hiring Remotely in United States
Internship
Internship SRE role responsible for availability, performance, and scalability of an e-commerce supply-chain platform. Tasks include SLO/SLA definition, observability (Prometheus/Grafana/Loki/Tempo/OpenTelemetry), incident response, capacity planning, disaster recovery for PostgreSQL, infrastructure-as-code (Terraform), CI/CD automation, and operational reliability for AI agent services. Mentored by Head of Technology/CTO with potential conversion to full-time based on performance.
The summary above was generated by AI
Overview: Reliability and Operational Excellence

The Site Reliability Engineer (SRE) is responsible for the ultimate stability, performance, and scalability of our entire integrated supply chain e-commerce platform. You will apply software engineering principles to operations, ensuring the high availability and resilience of the customer-facing e-commerce storefront, internal SaaS tools (WMS, OMS), and specialized AI agent services.

Internship Details

Duration: 3 months
Start Date: Immediate
Location: Remote
Stipend: None initially. Based on your first-quarter performance, you may be offered a paid full-time opportunity, or even be absorbed directly by the client as an FTE.

Key Responsibilities & Core Projects

You will be the champion of uptime, performance, and automated operations for systems handling the critical MES → WMS → OMS flow.

  • Availability & SLO Management: Define, implement, and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for core business processes and all application layers. Manage the platform's overall Service Level Agreement (SLA).

  • Observability & Alerting: Architect, maintain, and optimize the comprehensive observability stack (Prometheus, Grafana, Loki, Tempo, OpenTelemetry). Develop high-fidelity alerting and ensure distributed tracing across the NestJS modular monolith and associated data stores (PostgreSQL, Redis).

  • Incident Response & Review: Own the incident response workflow, ensuring rapid triage, mitigation, and root cause analysis. Conduct thorough post-incident reviews to drive continuous improvement and eliminate recurring toil.

  • Scalability & Capacity Planning: Optimize auto-scaling policies for all services running on Docker containers. Conduct capacity planning based on business projections, especially for peak e-commerce and manufacturing load.

  • Disaster Recovery (DR): Design, implement, and regularly test Disaster Recovery procedures, including backup and restoration workflows for PostgreSQL 15 using tools like pgBackRest.

  • Automation: Eliminate operational toil through automation, managing infrastructure-as-code (Terraform) and CI/CD pipelines (Makefile).

Required Technologies & Tools

Candidates must possess deep experience in cloud operations, observability, and infrastructure automation:

  • Observability Stack: Prometheus, Grafana, Loki, Tempo, OpenTelemetry (mandatory).

  • Infrastructure & Platform: Terraform, Docker, Traefik, Oracle Cloud Free VMs (or equivalent public cloud).

  • Data & Resilience: PostgreSQL (Deep knowledge), Redis, pgBackRest.

  • Automation: Strong scripting skills (Python/Bash) and experience with CI/CD tools and Makefile.

  • Methodology: Expert knowledge of SRE principles, toil reduction, and error budgeting.

AI Agent Focus

You will be responsible for the operational reliability of the emerging AI layer.

  • Agent Reliability: Implement specialized monitoring and logging for the AI agent services, ensuring LLM integrations and multi-agent systems (built with frameworks like LangChain) meet defined performance and availability SLOs.

  • Resource Optimization: Efficiently manage resource allocation for computationally intensive AI workloads to maintain platform stability and cost-efficiency.

Success Metrics & Career Path

Performance will be measured by:

  • Uptime/Availability: Achieving defined SLAs/SLOs across the platform.

  • MTTR: Reduction in Mean Time To Recover from production incidents.

  • Toil Reduction: Measured percentage reduction in manual, repetitive operational tasks through automation.

Mentorship Structure: Reports to the Head of Technology/CTO, working collaboratively with DevSecOps and development teams to ensure software is designed for reliability.

Similar Jobs

A Minute Ago
Remote or Hybrid
Senior level
Senior level
Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Manage tax compliance and advisory work, review tax returns, build client relationships, train staff, and participate in business development.
Top Skills: AccountingBusiness DevelopmentCpaTax AdvisoryTax Compliance
A Minute Ago
Remote or Hybrid
United States
129K-174K Annually
Senior level
129K-174K Annually
Senior level
Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
The Manager, Sales Architect leads the Sales Architect team, qualifying prospects, managing project handoffs, and designing solutions for technology sales.
Top Skills: AccountingEnterprise SoftwareErpInformation Systems
A Minute Ago
Remote or Hybrid
167K-213K Annually
Senior level
167K-213K Annually
Senior level
Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Manage tax compliance and advisory for high net worth clients, lead relationships, build teams, and ensure technical competency in tax services.
Top Skills: AccountingCpaCpeTax Compliance

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account