Filevine Logo

Filevine

Senior Site Reliability engineer

Posted 2 Days Ago
Remote
Hiring Remotely in USA
175K-195K Annually
Senior level
Remote
Hiring Remotely in USA
175K-195K Annually
Senior level
Design and improve observability (monitoring, logging, tracing, SLIs/SLOs), build automation and CI/CD, lead incident response and reliability improvements, mentor SREs, run on-call, and apply AI/ML to operational signals to forecast and reduce risks.
The summary above was generated by AI
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. Grounded in a singular system of truth, Filevine brings together data, documents, workflows, and teams into one unified platform—where modern legal work happens with clarity and consistency.
 
Powered by LOIS, the Legal Operating Intelligence System, Filevine connects context across every matter to transform legal operations from reactive to proactive. LOIS reads, understands, and reasons across your data to surface insight, automate complexity, and give professionals the clarity and confidence to see more, know more, and do more. Fueled by a team of exceptional collaborators and innovators, Filevine’s rapid growth has earned AI awards and recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.

Responsibilities

    • Design and improve the monitoring, logging, distributed tracing, dashboards, alerting, SLIs,
    and SLOs that give teams meaningful visibility into production health and customer impact.
    • Build and maintain automation, internal tools, and CI/CD systems that increase engineering
    efficiency, reduce toil, and support reliable deployments at scale. Take responsibility for the
    quality and reliability of tools and services you support.
    • Drive the implementation and continuous improvement of reliable systems for building,
    deploying, testing, and operating Filevine products, proactively identifying and resolving
    reliability, performance, scalability, and security risks before they impact customers.
    • Own complex production incidents through detection, triage, communication, resolution,
    and follow-up. Turn incident learning into durable corrective actions, stronger runbooks and
    operating practices, and improvements that reduce recurring incidents and operational
    burden.

    Lead significant technical initiatives from problem definition and design through
    implementation and adoption. Coordinate work across engineers and teams, communicate
    tradeoffs and risks, and help ensure the work delivers the intended results.
    • Mentor other Site Reliability Engineers through design reviews, incident follow-ups, paired
    problem-solving, and meaningful delegation. Help engineers develop stronger technical
    judgment and become increasingly capable of handling complex production work
    independently.
    • Participate in the shared on-call rotation and help ensure production systems are prepared
    to operate reliably at scale through capacity planning, operational readiness, and
    continuous improvements to resilience and recovery.
    • Apply AI and machine learning to analyze operational signals, identify patterns, forecast
    reliability and capacity risks, and implement improvements that make systems more
    reliable, efficient, and easier to operate.

Qualifications

    • 8+ years of hands-on experience in software engineering, cloud infrastructure, platform
    engineering, DevOps, or related technical roles, including at least 5 years in a Site Reliability
    Engineering or reliability-focused role.
    • Strong knowledge of distributed systems and hands-on experience operating Kubernetes
    workloads and cloud infrastructure in AWS or a comparable platform, with proficiency in
    Infrastructure as Code, monitoring, logging, alerting, distributed tracing, SLIs, and SLOs.
    • Strong proficiency with Python, Go, Bash, or a similar language, with demonstrated
    experience building and maintaining production tooling, automation, CI/CD pipelines, and
    deployment systems that reduce toil, improve reliability, and simplify ongoing operations.
    • Demonstrated ability to lead troubleshooting, incident response, root cause analysis, and
    long-term reliability improvements for complex production systems, including the
    elimination of recurring incidents and operational work.
    • Proven ability to mentor Site Reliability Engineers, help others build stronger technical
    judgment, communicate clearly with technical and business stakeholders, and lead complex
    initiatives from planning through delivery.
    • Demonstrated experience applying AI and machine learning to operational data and
    engineering workflows to identify patterns, forecast reliability or capacity risks, and
    implement measurable improvements with appropriate safeguards.

Cool Company Benefits:
- A dynamic, rapidly growing company, focused on helping organizations thrive 
- Medical, Dental, & Vision Insurance (for full-time employees)
- Competitive & Fair Pay
- Maternity & paternity leave (for full-time employees)
- Short & long-term disability
- Opportunity to learn from a dedicated leadership team
- Top-of-the-line company swag
 
Privacy Policy Notice
Filevine will handle your personal information according to what’s outlined in our Privacy Policy.
 
Communication about this opportunity, or any open role at Filevine, will only come from representatives with email addresses using "filevine.com". Other addresses reaching out are not affiliated with Filevine and should not be responded to.
 

Similar Jobs

5 Days Ago
In-Office or Remote
153K-205K Annually
Senior level
153K-205K Annually
Senior level
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate secure, scalable Kubernetes platforms and infrastructure as code (Terraform). Develop backend services and automation (Go, Python, JS/TS), improve CI/CD and observability, run on-call and incident response, define SLIs/SLOs and disaster recovery, embed security and compliance, mentor team members, and partner with product and engineering to raise reliability, performance, and cost-efficiency across hybrid and public-cloud environments.
Top Skills: Ci/CdGitopsGoJavaScriptKubernetesPythonTerraformTypescript
9 Days Ago
Remote or Hybrid
147K-278K Annually
Senior level
147K-278K Annually
Senior level
Cloud • Software
Design, deploy, and operate large-scale, multi-region cloud-native services to improve reliability, performance, and security. Partner with application teams to build automation, run SLO-driven incident response and on-call rotations, leverage Kubernetes and CNCF tooling, and implement scalable operations, chaos and scale testing, and infrastructure-as-code for a resilient SaaS platform.
Top Skills: ArgocdAWSGoKubernetesLinux/UnixOpentelemetryPrometheusPythonService Mesh
13 Days Ago
Remote
United States
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills: AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript

What you need to know about the Colorado Tech Scene

With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.

Key Facts About Colorado Tech

  • Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
  • Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
  • Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
  • Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
  • Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account