Metasys Logo

Metasys

DevOps Engineer Internship

Reposted 6 Days Ago
Remote
Hiring Remotely in United States
Internship
Remote
Hiring Remotely in United States
Internship
Build and maintain automated infrastructure and CI/CD pipelines using Terraform, Docker, Traefik, and Makefile. Implement observability (Prometheus, Grafana, Loki, Tempo, OpenTelemetry), backups (pgBackRest/Postgres15), and cloud/Linux administration. Automate deployment and monitoring for AI agent services and support monorepo workflow with SRE and DevSecOps teams.
The summary above was generated by AI
Overview: Infrastructure Automation and CI/CD

The DevOps Engineer is responsible for automating, streamlining, and maintaining the infrastructure and deployment pipelines for our entire integrated platform. You'll ensure rapid, reliable, and consistent delivery of our e-commerce storefront, internal supply chain tools (MES, WMS, OMS), and cutting-edge AI agent services, primarily utilizing Infrastructure-as-Code (IaC) and robust CI/CD practices.

Internship Details

Duration: 3 months
Start Date: Immediate
Location: Remote
Stipend: None initially. Based on your first-quarter performance, you may be offered a paid full-time opportunity, or even be absorbed directly by the client as an FTE.

Key Responsibilities & Core Projects

You will build and maintain the fully automated platform that underpins our entire tech stack.

  • Infrastructure-as-Code (IaC): Design, implement, and manage infrastructure provisioning across all environments using Terraform for our Oracle Cloud Free VMs (or equivalent cloud resources). Ensure infrastructure is auditable, repeatable, and secure.

  • CI/CD Pipeline Management: Set up and maintain the Continuous Integration and Continuous Deployment (CI/CD) pipelines, primarily driven by Makefile and automated testing, for the Node.js/NestJS modular monolith and Next.js frontend applications.

  • Containerization & Orchestration: Manage application containerization using Docker. Define deployment strategies, service discovery, and traffic routing using Traefik for our containerized services.

  • Observability Implementation: Implement, manage, and optimize the comprehensive logging, monitoring, and alerting system using our selected stack: Prometheus, Grafana, Loki, Tempo, and OpenTelemetry. Ensure end-to-end tracing is functional across the complex business flow (MES → WMS → OMS).

  • Resilience & Backups: Collaborate with the SRE team to implement high-availability features and maintain automated backup solutions, including pgBackRest for our PostgreSQL 15 database.

  • Workflow: Maintain the Monorepo structure for streamlined code management and deployment separation across applications (web / admin / API) and domain packages.

Required Technologies & Tools

Candidates must possess mandatory expertise in our core infrastructure and automation stack:

  • Infrastructure-as-Code: Expert proficiency in Terraform.

  • Containerization: Expert proficiency in Docker and deployment strategies (e.g., Traefik, orchestration concepts).

  • CI/CD: Hands-on experience building and maintaining complex pipelines (Makefile, Jenkins/GitHub Actions/GitLab CI concepts).

  • Observability: Strong implementation experience with Prometheus, Grafana, Loki, and OpenTelemetry.

  • Cloud & Linux: Experience with Linux administration and managing cloud resources (Oracle Cloud or equivalent).

AI Agent Focus

You will ensure the scalable and monitored deployment of the AI layer.

  • Deployment Automation: Automate the packaging and deployment pipelines for resource-intensive AI agent services and LLM fine-tuning environments.

  • Resource Monitoring: Set up specific monitoring and alerts to track the performance, resource consumption, and cost of the AI agent compute demands.

Success Metrics & Career Path

Performance will be measured by:

  • Deployment Frequency: Reduction in lead time and increased frequency of stable deployments.

  • Infrastructure Stability: Reliability of provisioned infrastructure (minimal unplanned downtime).

  • Observability Coverage: Completeness and reliability of monitoring, logging, and tracing across all production services.

Mentorship Structure: Reports to the Solution Architect or Head of Technology, working closely with the SRE, DevSecOps, and Backend engineering teams to build a robust platform.

Similar Jobs

4 Minutes Ago
Remote
United States
200K-210K Annually
Expert/Leader
200K-210K Annually
Expert/Leader
Information Technology • Software • Cybersecurity
Leads and scales the global Sales Engineering team while serving as Field CTO and the company’s external technical voice. Owns technical sales methodology, demos, evaluations, proof-of-value cycles, solution design, deployment handoffs, executive engagements, competitive positioning, market content, and field-driven product feedback. The role is a player-coach position requiring deep cybersecurity expertise, executive presence, public speaking, and significant travel.
Top Skills: AIAstEndpoint SecurityMalware AnalysisOt/IcsPacket CaptureScaSoc OperationsSoftware Supply Chain SecuritySpectraThreat Intelligence
7 Minutes Ago
Remote or Hybrid
120K-150K Annually
Senior level
120K-150K Annually
Senior level
Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Leads automotive functional safety development and reviews under ISO 26262. Develops safety concepts, architectures, requirements, analyses, safety cases, and compliance assessments for embedded software, hardware, and vehicle systems. Advises customers, leads workshops and training, conducts technical reviews and gap assessments, and supports safety evaluations across ADAS, electrified vehicles, battery systems, and other safety-critical technologies. Limited client travel may be required.
Top Skills: AdasAutomated Driving SystemsAutomotive SoftwareBattery Management SystemsDfaElectronic SystemsEmbedded SoftwareFmeaFtaHaraIec 62508Iso 21434Iso 21448Iso 26262Iso 42001Iso/Pas 8800Power ElectronicsStpa
8 Minutes Ago
Remote or Hybrid
United States
Senior level
Senior level
Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Oversees safety, durability, and regulatory testing of toy products. Examines samples, operates and maintains laboratory equipment, builds test setups, evaluates results, prepares reports, and recommends solutions to testing issues. Trains laboratory staff, communicates with clients and engineering teams, supports test method development, and incorporates automation and continuous improvement into testing processes.
Top Skills: Laboratory Testing EquipmentMS OfficeTest Automation

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account