Senior Site Reliability Engineer (SRE)
TechnologieDescription
About AlayaCare
At AlayaCare, we are more than a rapidly growing SaaS company: we are a passionate team transforming home care. Our cloud platform enables care providers worldwide to deliver better outcomes to their clients.
With over 550 employees across Canada, the United States, Australia, and Brazil, we are united by a common mission and a strong culture of transparency, growth, and human connection. Whether you're just starting out or a seasoned expert, AlayaCare offers you the opportunity to grow your impact, skills, and career.
About the Role
We are seeking a Senior Site Reliability Engineer to join our SRE team. You will be responsible for scaling AWS cloud infrastructure, evolving Kubernetes deployment pipelines, improving monitoring, alerts, and resilience, and developing tools enabling product teams to deliver safely and effectively. For acquired products based on Azure, the focus is on monitoring, alert triage, and incident management guided by runbooks, rather than designing platforms from scratch.
This role is responsible for shared platform services across cloud regions, including databases, messaging, logging, search, and tenant provisioning. The Senior Site Reliability Engineer is called upon to lead major infrastructure initiatives and proof-of-concept projects, contribute to technical planning and prioritization, and collaborate with Product teams to reduce operational incidents.
The SRE team also develops and operates AI-based tools to optimize runbooks, accelerate incident response, and generate operational insights from platform telemetry, aiming to improve reliability and reduce manual effort.
Development, Automation, and Tooling
- Design, build, and maintain infrastructure and platform services, including Kubernetes and observability tools.
- Implement infrastructure as code, configuration management, and automated testing to ensure reliable and reproducible environments.
- Contribute to code and configuration reviews to improve scalability, maintainability, and reusability.
Reliability and Operations
- Monitor production systems, troubleshoot issues, and improve logging, monitoring, alerts, and runbooks.
- Participate in on-call and support rotations, incident management, and post-incident reviews to improve long-term reliability.
Requirements and Collaboration
- Collaborate with Product, Engineering, and Development teams to translate requirements into reliable and operable infrastructure solutions.
- Identify risks related to operability, security, performance, and costs, and recommend pragmatic trade-offs.
Continuous Improvement
- Contribute to operational quality through runbooks, security hardening, performance optimization, and process improvement.
- Stay up-to-date on emerging SRE practices, including AI-assisted operations and modern AWS platform patterns.
What You Bring to the Team
- Bachelor's degree or higher in computer science, computer engineering, or a related field, with demonstrated practical experience.
- 5+ years of practical experience.
- Strong practical experience with AWS in a multi-account and multi-region environment: EKS, AWS Organizations, IAM, and KMS.
- Advanced proficiency in Terraform and Infrastructure as Code workflows, including Atlantis/GitOps, state management, and module/provider upgrades.
- Practical experience running workloads on Docker and Kubernetes in production, including Gateway API ingress patterns and cluster lifecycle management (upgrades, extensions, node provisioning).
- Strong experience in Linux system administration and production troubleshooting.
- Proficiency in at least one development or scripting language such as Python, Go, or Bash.
- Experience with an observability platform (e.g., New Relic, OpenSearch, CloudWatch, OpenTelemetry) and event-driven alerts (e.g., EventBridge, SNS, PagerDuty).
- Knowledge of system and network security fundamentals, including WAF, least-privilege IAM, secrets management, and disaster recovery.
- Experience in incident management (on-call, triage, remediation, post-incident review) and writing operational runbooks.
- Practical experience operating Aurora MySQL and PostgreSQL in production, including migrations, performance tuning, and backup/restore.
- Excellent communication and collaboration skills, with the ability to work effectively with technical and non-technical teams in a distributed environment.
- Experience in cloud cost optimization (rightsizing, reserved capacity, cost tagging) is a plus.
- Experience with Flux/ArgoCD, Karpenter, or GPU infrastructure Ray/Anyscale is a plus.
- Experience in monitoring production systems on Azure is a plus.
- Cloud or Kubernetes certifications are a plus.
- Bilingual French-English is a plus.
- Why Join AlayaCare?
- Give Your Work Meaning
At AlayaCare, you will help design technology that empowers care providers and improves outcomes for patients and their families. Every line of code and every interaction with a client contributes to making care more connected, accessible, and human.
Grow in a Culture of Trust
We believe in transparency, feedback, and assuming good intentions. Here, you will feel safe to share your ideas and career goals. You will be supported in achieving them through mentorship, career mobility, and our internal promotion philosophy.
A Work-Life Balance That Works for You
We prioritize flexibility and well-being. From Wellness Fridays to volunteer time off, passing by flexible vacation, we ensure you have space to recharge, contribute to your community, and live your best life.
Benefits That Matter
- Stock options in a growing, well-funded company.
- Comprehensive health benefits, telemedicine, and lifestyle spending accounts.
- Parental leave supplement and family support programs.
- Inclusive Design
We celebrate diverse perspectives and foster belonging through our DEIB (Diversity, Equity, Inclusion, and Belonging) initiatives. Employee-led events, summits, and social activities, both in-person and virtual, create meaningful connections among our global teams.
Work Location and Work Model
This position is located in the Greater Montreal Area. At AlayaCare, our hybrid model includes days of in-office collaboration, and team members are expected to be present on those days to foster connections, innovation, and teamwork.
Ready to Join Us?
Apply today and become part of a company that makes a real difference in the future of home and community care. Not the right fit for you? Share this offer with someone who might be an excellent fit.
AlayaCare uses AI tools during its hiring process to support fair, consistent, and objective decision-making. Some initial screening steps may be automated to help identify qualified candidates. If your application is automatically rejected, you may request a human review.
We are committed to creating a workplace where everyone and everything belongs.
Cette offre a été agrégée depuis indeed. Groupe Sentinella n'est pas l'employeur; postuler vous redirige vers le site original. Le texte intégral appartient à l'auteur de l'offre.
Envie qu'on travaille pour vous ?
Inscrivez-vous au banc de talents Sentinella. On vous contacte dès qu'un mandat correspond à votre profil — y compris des postes comme celui-ci.
Rejoindre le banc de talents