Fleetio Logo

Fleetio

Senior Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Scale and improve the reliability, performance, and scalability of Fleetio’s Ruby on Rails application and infrastructure. Responsibilities include observability, incident response, AI-powered automation, database capacity planning, disaster recovery, backups, documentation, performance analysis, code review, and on-call participation. The role collaborates with engineering teams to address bottlenecks and adopt AI-driven reliability practices.
The summary above was generated by AI

A little about us…Fleetio is a modern software platform that helps thousands of organizations worldwide manage their fleet operations. Transportation technology is a hot market, and we’re leading the charge with raving fans and new customers signing up every day. We raised $450M in our Series D funding round in March of 2025 and are on an exciting trajectory as a company. Fleetio is also a proud founding member of the Rails Foundation!

More about our team and company:

  • Fleetio overview video: https://www.youtube.com/watch?v=YoXyXTFWbkg
  • Our careers page: https://www.fleetio.com/careers
Description

Our Platform Engineering team is looking for a Senior Site Reliability Engineer to help run, maintain, and improve the performance of our Ruby on Rails Stack and Infrastructure. You will help scale our application using best in class architecture and software design. This includes training, software engineering, system design, and operational practices that support the needs of our engineers and customers while accounting for future growth. You will be entrusted with proactively identifying and owning initiatives that will help improve the performance, reliability, and scalability of our application stack and databases.

Our team treats AI as a core part of how we engineer. We build and use AI agents, skills, and automated workflows to handle toil, speed up investigation and remediation, and give our engineers more time for high-leverage work.

More About Our Team and Company
  • Watch our culture videos: https://fleet.io/culture
  • Engineering culture, interview process and videos: https://www.fleetio.com/careers/engineering
  • Fleetio Go overview video: https://www.fleetio.com/go
  • More about the Fleetio platform: https://www.fleetio.com/features
  • API docs: developer.fleetio.com
  • Test drive Fleetio to get an even better feel for what we're building: https://www.fleetio.com/register

This is a remote opportunity and is open to candidates in the United States.

Who You Are

Our ideal candidate is an Infrastructure Engineer experienced in scaling Ruby on Rails applications, with a passion for optimization and performance improvements. You bring a strong background in Site Reliability and Infrastructure Engineering for Rails applications. You follow Agile and DevOps principles, can effectively influence teams to achieve goals, and demonstrate excellent problem-solving skills in our fast-paced environment.

You're curious about how AI is changing infrastructure and reliability work, and you're eager to shape how a platform team puts it to use.

Your Impact

As a Senior Site Reliability Engineer on Fleetio's Platform Engineering team, you will:

  • Proactively identify, triage, and resolve performance issues
  • Enhance system observability by monitoring performance metrics across Ruby, Rails, and database systems, including SLOs and SLIs
  • Build and maintain AI agents, skills, and automations that reduce operational toil across incident response, triage, and routine maintenance
  • Use AI-assisted tooling to accelerate performance analysis, root-cause investigation, and code review
  • Collaborate with other SREs to proactively identify and address performance bottlenecks
  • Help product engineers adopt AI-driven workflows for performance and reliability best practices
  • Lead database capacity planning and upgrade initiatives
  • Manage the database-specific components of disaster recovery planning and execution
  • Oversee backup systems and pre-production databases
  • Create and maintain infrastructure and operations documentation, including runbooks and context that both engineers and AI agents can act on
  • Participate in the on-call rotation
Your Experience
  • 5+ years of Ruby/Rails Experience
  • 3+ years of AWS Experience
  • Kubernetes experience
  • Experience with profiling and benchmarking source code
  • Effective at code review and identifying potential performance problems before they reach production
  • Experience with Datadog or other APM tools
  • Excellent written and verbal communication skills
Considered a Plus
  • Experience building AI agents, LLM-powered automations, or integrations (e.g., using tool/function calling or MCP) for engineering or operations workflows
  • Infrastructure as Code tools (Terraform)
  • Deep understanding of cloud network fundamentals (routing, firewalls, load balancers, CDNs, VPCs, etc.)
  • Experience with distributed event and data stores, such as Kafka, Redis, Elasticsearch, Memcached, and TimescaleDB
  • You know a thing or two about the fleet management industry
Benefits 
  • Multiple health/dental coverage options (100% coverage for employee, 50% for family)
  • Vision insurance
  • Incentive stock options
  • 401(k) match of 4%
  • PTO - 4 weeks (increases at year two!)
  • 12 company holidays + 2 floating holidays
  • Parental leave - birthing parent (16 weeks paid) non-birthing (4 weeks paid)
  • FSA & HSA options
  • Short and long term disability (short term 100% paid)
  • Community service funds
  • Professional development funds
  • Wellbeing fund - $150 quarterly
  • Business expense stipend - $125 quarterly
  • Mac laptop + new hire equipment stipend
  • Fully stocked kitchen with tons of drinks & snacks (BHM only)
  • Remote working friendly since 2012 #LI-Remote

Fleetio provides equal employment opportunities to all employees and applicants and prohibits discrimination and harassment. We celebrate diversity and are committed to creating an inclusive environment for all. All employment is decided on the basis of qualifications, merit and business need.

This application is not intended to and does not create a contract or offer of employment. Employment with Fleetio is at will.

If you have a disability or a special need that requires an accommodation to fill out the online application, please let us know by calling (205) 718-7500.

Similar Jobs

8 Days Ago
In-Office or Remote
75K-195K Annually
Senior level
75K-195K Annually
Senior level
Cloud • Software
Own NetBox’s build and release pipeline from image creation through Cloud and Enterprise deployment. Improve Django and PostgreSQL performance, establish observability and SLOs, strengthen software supply chain security, and support SOC 2 compliance. Participate in on-call rotations, incident response, and postmortems. Drive cross-team migrations and release processes while contributing fixes to NetBox Core when reliability issues originate in the application.
Top Skills: ArgocdAws Ec2Aws IamAws RdsAws VpcClaude CodeCosignDjangoFluxcdGithub ActionsGrafanaHelmKubernetesPostgresPrometheusPythonSigstoreSlsaTerraform
21 Days Ago
Easy Apply
Remote or Hybrid
USA
Easy Apply
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Healthtech • Information Technology • Software • Telehealth
Develop, monitor, and maintain distributed production systems and AWS-based microservices infrastructure. Build automation, tooling, and repeatable processes that improve uptime, scalability, security, and operational efficiency. Support product engineering teams with performance, scaling, incident diagnosis, and production debugging. Analyze and tune systems, code, and networking while participating in on-call operations and blameless post-mortems.
Top Skills: AWSDnsDockerGCPGenaiHttp/HttpsKubernetesLoad BalancersNtpReverse ProxiesTcp/IpTlsWeb Application Firewalls
22 Days Ago
Remote
United States
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Own reliability, scalability, security, observability, and incident response for production applications across AWS and on-premises DoD environments. Build monitoring and alerting, define SLIs and SLOs, lead post-incident reviews, automate infrastructure with Terraform and Ansible, operate Kubernetes clusters, embed RMF and STIG controls, reduce operational toil, and support secure air-gapped deployments.
Top Skills: AlloyAnsibleAWSAws GovcloudBashDatadogElk StackGithub ActionsGitlab Ci/CdGitopsGoGrafanaHyper-VIstioJenkinsKubernetesLinkerdLokiNutanixPrometheusProxmoxPythonRmfSecurity+StigsTerraformVMware

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account