Lead observability and incident management efforts: define SLIs/SLOs, build monitoring/alerting, dashboards, logging, and tracing. Drive incident response, postmortems, and reliability improvements to reduce MTTD/MTTR. Integrate observability into CI/CD, maintain AWS and Kubernetes infrastructure, automate operations, and mentor engineers on SRE best practices.
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. Grounded in a singular system of truth, Filevine brings together data, documents, workflows, and teams into one unified platform—where modern legal work happens with clarity and consistency.
Powered by LOIS, the Legal Operating Intelligence System, Filevine connects context across every matter to transform legal operations from reactive to proactive. LOIS reads, understands, and reasons across your data to surface insight, automate complexity, and give professionals the clarity and confidence to see more, know more, and do more. Fueled by a team of exceptional collaborators and innovators, Filevine’s rapid growth has earned AI awards and recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Role Summary
As a Site Reliability Engineer at Filevine, you will improve the reliability, scalability, and
operational maturity of the Filevine platform. You’ll design automation that reduces toil,
strengthen observability, support reliable deployments at scale, and solve production challenges
that keep Filevine running for legal teams across the country. This role is built for engineers
who apply software engineering principles to infrastructure problems, thrive on complex
technical challenges, and are energized by taking ownership of the systems they build and
operate — growing into deeper expertise as their platform knowledge expands
As a Site Reliability Engineer at Filevine, you will improve the reliability, scalability, and
operational maturity of the Filevine platform. You’ll design automation that reduces toil,
strengthen observability, support reliable deployments at scale, and solve production challenges
that keep Filevine running for legal teams across the country. This role is built for engineers
who apply software engineering principles to infrastructure problems, thrive on complex
technical challenges, and are energized by taking ownership of the systems they build and
operate — growing into deeper expertise as their platform knowledge expands
Responsibilities
Design, build, and maintain the monitoring, logging, distributed tracing, dashboards, and
alerting that give teams meaningful visibility into production health.
Build automation, tooling, and CI/CD improvements that increase engineering efficiency,
reduce toil, and support reliable deployments at scale.
Design, implement, and maintain reliable systems for building, deploying, testing, and
operating Filevine products — proactively identifying and resolving reliability,
performance, scalability, and security risks before they impact customers.
Participate in a shared 24/7 on-call rotation, using operational insights to drive automation
and long-term reliability improvements; continuously improve runbooks, documentation,
and engineering standards.
Take ownership of technical initiatives from design through implementation, develop deep
expertise in critical areas of the Filevine platform, and communicate clearly with technical
and business stakeholders.
alerting that give teams meaningful visibility into production health.
Build automation, tooling, and CI/CD improvements that increase engineering efficiency,
reduce toil, and support reliable deployments at scale.
Design, implement, and maintain reliable systems for building, deploying, testing, and
operating Filevine products — proactively identifying and resolving reliability,
performance, scalability, and security risks before they impact customers.
Participate in a shared 24/7 on-call rotation, using operational insights to drive automation
and long-term reliability improvements; continuously improve runbooks, documentation,
and engineering standards.
Take ownership of technical initiatives from design through implementation, develop deep
expertise in critical areas of the Filevine platform, and communicate clearly with technical
and business stakeholders.
What we are looking for
4+ years of hands-on experience in software engineering, cloud infrastructure, platform
engineering, DevOps, or related technical roles, including at least 2 years in a Site Reliability
Engineering or reliability-focused role.
-Working knowledge of distributed systems and how applications, infrastructure, and cloud
services interact in production; demonstrated ability to troubleshoot production issues,
perform root cause analysis, and drive long-term reliability improvements.
-Proficiency with Python, Bash, or similar scripting languages; experience building
production tooling, automation, or CI/CD pipelines and deployment automation.
-Hands-on experience operating Kubernetes-based workloads and cloud infrastructure in
AWS or a comparable platform, including compute, container orchestration, networking,
IAM, object storage, and cloud-native monitoring.
-Experience with Infrastructure as Code tools such as Terraform, Pulumi, or AWS
CloudFormation, and familiarity with modern observability practices including monitoring,
logging, alerting, distributed tracing, and incident response.
-Experience using AI-assisted engineering tools to improve productivity, accelerate
troubleshooting, or automate operational tasks; curiosity, ownership, and a passion for
building reliable systems through continuous improvement.
-Strong written and verbal communication skills; Bachelor’s degree in Computer Science,
Information Systems, or a related field, equivalent industry certifications, or comparable
professional experience.
engineering, DevOps, or related technical roles, including at least 2 years in a Site Reliability
Engineering or reliability-focused role.
-Working knowledge of distributed systems and how applications, infrastructure, and cloud
services interact in production; demonstrated ability to troubleshoot production issues,
perform root cause analysis, and drive long-term reliability improvements.
-Proficiency with Python, Bash, or similar scripting languages; experience building
production tooling, automation, or CI/CD pipelines and deployment automation.
-Hands-on experience operating Kubernetes-based workloads and cloud infrastructure in
AWS or a comparable platform, including compute, container orchestration, networking,
IAM, object storage, and cloud-native monitoring.
-Experience with Infrastructure as Code tools such as Terraform, Pulumi, or AWS
CloudFormation, and familiarity with modern observability practices including monitoring,
logging, alerting, distributed tracing, and incident response.
-Experience using AI-assisted engineering tools to improve productivity, accelerate
troubleshooting, or automate operational tasks; curiosity, ownership, and a passion for
building reliable systems through continuous improvement.
-Strong written and verbal communication skills; Bachelor’s degree in Computer Science,
Information Systems, or a related field, equivalent industry certifications, or comparable
professional experience.
Cool Company Benefits:
- A dynamic, rapidly growing company, focused on helping organizations thrive
- Medical, Dental, & Vision Insurance (for full-time employees)
- Competitive & Fair Pay
- Maternity & paternity leave (for full-time employees)
- Short & long-term disability
- Opportunity to learn from a dedicated leadership team
- Top-of-the-line company swag
Privacy Policy Notice
Filevine will handle your personal information according to what’s outlined in our Privacy Policy.
Communication about this opportunity, or any open role at Filevine, will only come from representatives with email addresses using "filevine.com". Other addresses reaching out are not affiliated with Filevine and should not be responded to.
Similar Jobs
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
Design, build, and operate scalable blockchain infrastructure and Kubernetes platforms. Implement IaC, CI/CD, AI-powered automation, monitoring, incident response, and reliability improvements. Mentor engineers, lead cross-functional initiatives, and support network launches, upgrades, and production troubleshooting in a follow-the-sun on-call rotation.
Top Skills:
Agentic AutomationArcBaseBlue-Green DeploymentCanary ReleasesChaos EngineeringCi/CdCloud-Native ToolingContainerizationControllersDnsEthereumGenerative AiGoHelmInfrastructure As CodeKubernetesLoad BalancersMcp ServersObservability ToolingOperatorsPulumiPythonRbacSolanaSQLTerraformVpc
Financial Services
Lead SRE responsible for driving reliability, resiliency design reviews, major-incident leadership, mentoring engineers, establishing SLOs/error budgets, improving observability, CI/CD and container practices, and adopting AI-assisted reliability workflows with appropriate guardrails.
Top Skills:
.NetCi/CdContainer OrchestrationContainersEnterprise AiJavaMonitoringNetworkingObservabilityPythonSpring BootTelemetry
Financial Services
Lead SRE responsible for resiliency design reviews, mentoring, SRE best-practice adoption, building IaC and CI/CD pipelines, operating containerized services, observability and SLO-driven incident prevention, 24x7 production support, and driving AI-assisted reliability workflows with governance and auditability.
Top Skills:
.NetAWSCi/CdDatadogDnsDockerDynatraceEcsGitlabGrafanaJavaJenkinsKafkaKubernetesLinuxLoad BalancingPrometheusPythonSplunkSpring BootTcp/IpTerraformTls
What you need to know about the Colorado Tech Scene
With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.
Key Facts About Colorado Tech
- Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
- Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
- Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
- Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
- Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute


