The Senior Site Reliability Engineer improves production reliability, resilience, and system availability. Responsibilities include operating Linux and Kubernetes infrastructure, automating workflows, developing tools, managing observability, responding to incidents, participating in on-call rotations, writing runbooks, and conducting blameless postmortems. The role collaborates with distributed teams and stakeholders, supports cloud infrastructure, troubleshoots live systems, and promotes DevOps and SRE best practices.
About Megaport
We’re not your typical tech company – and we don’t want to be. Megaport is the global leader in Network as a Service (NaaS), and has transformed the way businesses connect to the cloud, data centers, and each other. We’re publicly listed on the Australian Stock Exchange and partnered with the biggest names in tech like Amazon, Microsoft, Google, Oracle, IBM, and more. Headquartered in Brisbane with a crew of over 600 people spread across Asia-Pacific, Europe, and the Americas, our employees enjoy an environment that is collaborative, supportive, and (actually) fun.
Our Team Culture
We’re a team of problem solvers, pixel pushers, code slingers, and cloud fanatics. Culture is more than a poster on the wall here – collaboration beats hierarchy, curiosity fuels our growth, and everyone’s voice matters. We take our work seriously, but not ourselves. We work across time zones to execute on our global vision, trust each other to get things done, and never compromise our values for commercial gain. Most importantly, we place our customers at the center of everything we do.
We’re committed to increasing representation in the tech industry and welcome applicants from all backgrounds. Don’t meet every requirement? That’s okay. If you’re excited about this role, we encourage you to apply.
As a Senior Platform Engineer, you are a champion for DevOps and SRE culture and industry best practice within Megaport. You will work alongside talented team members in multiple timezones ensuring that systems are secure, maintainable and available. External to the team you will be engaging with stakeholders in requirements analysis and demonstrations. Technically you will be very hands on. Continually evolving your skills through a mix of peer reviews and research. Ultimately your obsession is customer success and ensuring company goals are met.
What You Will Be Doing
- Improving production reliability and system resilience within an SRE scoped team
- Championing high standards of work and industry best practices
- Communicating with teams and stakeholders at all stages
- Bringing fresh ideas to the table and encouraging others
- Diving into complex technical problems with a can-do attitude
- Working across numerous technologies in a fast-changing industry
- Participating in on-call rotation, incident response, and blameless post-incident reviews
- Writing code, handling alerts, improving solutions, and supporting others
- Playing a crucial role in the success of your company and team
What We Are Looking For
- 5+ years administering Linux systems and related infrastructure in production environments
- A collaborative SRE mindset, with familiarity around SLIs/SLOs/SLAs, error budgets, blast radius, and blameless postmortems
- A focus on automation, reducing toil, and preventing problem recurrence
- A track record of writing runbooks that work for the broader team, not just yourself
- Strong Kubernetes and broader ecosystem fundamentals
- Cloud infrastructure experience; AWS strongly preferred and bare-metal is a bonus
- Strong tool development - Bash, plus either Python or Go preferred, or similar
- Infrastructure-as-code tooling experience - Terraform preferred
- CI/CD and version control, GitHub preferred
- Database experience - one of Postgres, Cassandra, or ClickHouse preferred
- Experience operating a production observability stack (metrics, logs, traces), with an eye for signal over noise
- Comfortable working on live production infrastructure, with strong troubleshooting instincts and ownership of incident response
- A history of continual professional development
- A self-directed style suited to an async, globally distributed team, and comfortable picking up adjacent work when the situation calls for it
What We Offer
- Flexible working environment – a remote-first culture with coworking options available.
- Generous leave plans – including 4 weeks of paid annual leave, parental leave, birthday leave, and a purchased annual leave program.
- Health and wellness support – through a wellness allowance and employee wellbeing initiatives.
- Comprehensive learning support – generous study and training allowance plus 5 days of paid study leave
- Creative, modern workspaces – designed to inspire when you're not working remotely
- Motivated, inclusive team – work alongside industry experts and fresh talent
- Recognition programs – celebrate achievements with our Legend and Kudos awards
#LI-DNI
If you have any questions, please reach out to Megaport's Talent Acquisition Team at [email protected]
NOTE: All Megaport business correspondence is conducted via our business email accounts (@megaport.com). If you have any concerns, please reach out to Megaport's careers team [email protected] directly and we will verify the legitimacy of any communication. Megaport will not ask you to create an account via Microsoft teams, and does not associate with any email accounts under "@megaportau.com".
All applications will be treated in confidence.
Please see Part 2 of our Privacy Policy to see what information Megaport collects from job applicants, why, and how we store and use it. Note that you’re entitled to know what personal data of yours Megaport holds, to request updates, rectification, and in some circumstances restriction or deletion thereof if you object (you being entitled to withdraw your consent to our holding your information at any time). Please see Part 5 of our Privacy Policy for more details on this and how to contact Megaport's data protection officer if you have any further privacy-related questions. Candidates who meet the selection criteria will be invited to attend an interview. Strictly no Recruitment Agencies.
Similar Jobs
Healthtech • Information Technology • Software • Telehealth
Develop, monitor, and maintain distributed production systems and AWS-based microservices infrastructure. Build automation, tooling, and repeatable processes that improve uptime, scalability, security, and operational efficiency. Support product engineering teams with performance, scaling, incident diagnosis, and production debugging. Analyze and tune systems, code, and networking while participating in on-call operations and blameless post-mortems.
Top Skills:
AWSDnsDockerGCPGenaiHttp/HttpsKubernetesLoad BalancersNtpReverse ProxiesTcp/IpTlsWeb Application Firewalls
Software • Defense
Own reliability, scalability, security, observability, and incident response for production applications across AWS and on-premises DoD environments. Build monitoring and alerting, define SLIs and SLOs, lead post-incident reviews, automate infrastructure with Terraform and Ansible, operate Kubernetes clusters, embed RMF and STIG controls, reduce operational toil, and support secure air-gapped deployments.
Top Skills:
AlloyAnsibleAWSAws GovcloudBashDatadogElk StackGithub ActionsGitlab Ci/CdGitopsGoGrafanaHyper-VIstioJenkinsKubernetesLinkerdLokiNutanixPrometheusProxmoxPythonRmfSecurity+StigsTerraformVMware
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own the reliability, performance, resilience, observability, and security of AWS and Kubernetes infrastructure supporting products and AI/ML workloads. Define SLOs, lead incident response and root-cause analysis, build Terraform automation, optimize cloud costs, reduce operational toil, and establish deployment standards that help engineers ship reliably. Participate in on-call rotations and maintain HIPAA-compliant infrastructure.
Top Skills:
AWSClaudeDatadogGitlabGoHipaaIstioKubernetesNatsPostgresPythonSoc 2TerraformTypescript
What you need to know about the Colorado Tech Scene
With a business-friendly climate and research universities like CU Boulder and Colorado State, Colorado has made a name for itself as a startup ecosystem. The state boasts a skilled workforce and high quality of life thanks to its affordable housing, vibrant cultural scene and unparalleled opportunities for outdoor recreation. Colorado is also home to the National Renewable Energy Laboratory, helping cement its status as a hub for renewable energy innovation.
Key Facts About Colorado Tech
- Number of Tech Workers: 260,000; 8.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Lockheed Martin, Century Link, Comcast, BAE Systems, Level 3
- Key Industries: Software, artificial intelligence, aerospace, e-commerce, fintech, healthtech
- Funding Landscape: $4.9 billion in VC funding in 2024 (Pitchbook)
- Notable Investors: Access Venture Partners, Ridgeline Ventures, Techstars, Blackhorn Ventures
- Research Centers and Universities: Colorado School of Mines, University of Colorado Boulder, University of Denver, Colorado State University, Mesa Laboratory, Space Science Institute, National Center for Atmospheric Research, National Renewable Energy Laboratory, Gottlieb Institute



.png)