Site Reliability Engineer (SRE)
TechnologieDescription
About Us
Toboggan Labs is a boutique consultancy building at the intersection of AI and healthcare. We solve challenging human problems by applying cutting-edge technology and domain understanding.
About the role
We're seeking a Site Reliability Engineer (SRE) to help our clients build reliable, observable, and secure production systems.
In this role, you will work closely with client engineering and operations teams to improve system reliability, reduce toil, and build the operational foundations — deployment pipelines, monitoring, incident management, and infrastructure — that keep production systems running smoothly.
Note that while we specialize in healthcare and regulated industries, not all our projects are in these fields, so you may work across different domains from time to time.
Your work will consist of
- Infrastructure and security engineering — Design and maintain resilient, secure cloud infrastructure using infrastructure-as-code; implement security controls, hardening standards, and compliance guardrails across client environments.
- Observability and reliability — Design and implement monitoring, alerting, and logging systems; lead incident response and post-mortem processes; define and track SLOs and SLIs.
- Automation and platform operations — Automate deployment pipelines, infrastructure provisioning, and operational runbooks to reduce toil and improve system resilience.
- On some projects, technical leadership — Own the reliability and infrastructure workstream, guide client engineering teams on SRE practices, and contribute to architectural decisions.
- Supporting the team — Share SRE expertise with colleagues, contribute to internal tooling and documentation, mentor team members, and participate in the broader Toboggan community.
About you
We are seeking individuals with a strong background in cloud infrastructure, DevOps, and reliability engineering, who bring security mindedness to everything they build. Most of our clients run AWS, Terraform, GitHub Actions or similar CI/CD tooling. You should be comfortable working at the intersection of reliability, security, and IT operations.
When we say SRE we mean someone who treats reliability as a feature, not an afterthought — someone who writes code to eliminate toil, thinks in systems, and builds infrastructure that is as secure as it is observable.
Cette offre a été agrégée depuis indeed. Groupe Sentinella n'est pas l'employeur; postuler vous redirige vers le site original. Le texte intégral appartient à l'auteur de l'offre.
Envie qu'on travaille pour vous ?
Inscrivez-vous au banc de talents Sentinella. On vous contacte dès qu'un mandat correspond à votre profil — y compris des postes comme celui-ci.
Rejoindre le banc de talents