Job Description
Your safety comes first — read before you go. Taf4All only lists job offers published by third parties. We are not the employer, we do not conduct these recruitments, and we cannot guarantee what happens once you make contact — you deal directly with the person or company behind the offer, at your own risk.
Our client’s Cloud Operations team is growing its Site Reliability Engineering (SRE) function to enhance the reliability and performance of user-facing services and production systems. This role demands a unique combination of hands-on operational expertise and software development proficiency, where engineers apply robust engineering principles, operational rigor, and automation to both infrastructure and code. The team focuses on core systems domains—including networking, Linux kernel optimization, scaling strategies, algorithmic efficiency, and distributed system architectures. In this position, you will participate in an on-call rotation to address production availability incidents, collaborate with service engineers to resolve customer-facing issues, and proactively leverage on-call time to mitigate potential incidents before they occur. Responsibilities include managing infrastructure through tools like Ansible, Puppet, Terraform, and Kubernetes, ensuring monitoring and alerting systems are symptom-based rather than outage-driven. You will document all actions meticulously to transform findings into repeatable processes—and ultimately, into automated solutions—while refining deployment workflows to achieve seamless, predictable releases. Additionally, you will design, construct, and maintain scalable infrastructure capable of supporting hundreds of thousands of concurrent users, troubleshoot complex production issues across service layers, and strategically plan infrastructure expansion. Ideal candidates will demonstrate a cloud-first mindset, regardless of the underlying public cloud platform, and prioritize security in all initiatives. They will possess deep technical curiosity about systems—analyzing edge cases, failure modes, and implementation-specific behaviors—alongside strong proficiency in Linux and Windows environments. Familiarity with configuration management tools such as Ansible or Puppet, as well as programming expertise in languages like Python, Java, Go, or Node.js, is essential. The ability to work asynchronously, communicate clearly through documentation, and adopt a proactive “fix-it-now” attitude is critical.
Expérience
Experience with technologies such as Nginx, HAProxy, Docker, Kubernetes, Terraform, or equivalent platforms is highly valued. Potential projects include developing infrastructure automation with Ansible and Terraform, enhancing Prometheus monitoring or designing new metrics, assisting in software deployments, and leading the migration from traditional AWS virtual machines to cloud-native, containerized deployments on Kubernetes (EKS). You may also collaborate with product teams to define SRE key performance indicators, as this SRE function is in its early stages of development.
Formation / Diplômes
Secteur : Engineering and Information Technology
---
**
[Click the Apply button below to apply, and Create my CV to build a CV tailored to this offer, professionally]
Ready to apply?
Get seen by recruiters
Publish your CV on the Candidates page — the place recruiters browse directly to find profiles like yours.
See the Candidates pageYou are recruiting?
Find candidates on TAF4ALL🚀 Boost your application
Stand out with a professional CV and a personalized cover letter generated by AI in 3 minutes. 3 Ingénieur position(s) posted this week in Madagascar.
Canadian employers are also recruiting in Africa
Real offers from Canadian employers who are explicitly looking for candidates outside Canada. No agency, no fees — you apply yourself.
Expert Application Advice
Our client’s Cloud Operations team is growing its Site Reliability Engineering (SRE) function to enhance the reliability and performance of user-facing services and production systems. This role demands a unique combination of hands-on operational expertise and software development proficiency, where engineers apply robust engineering principles, operational rigor, and automation to both infrastructure and code. The team focuses on core systems domains—including networking, Linux kernel optimization, scaling strategies, algorithmic efficiency, and distributed system architectures. In this position, you will participate in an on-call rotation to address production availability incidents, collaborate with service engineers to resolve customer-facing issues, and proactively leverage on-call time to mitigate potential incidents before they occur. Responsibilities include managing infrastructure through tools like Ansible, Puppet, Terraform, and Kubernetes, ensuring monitoring and alerting systems are symptom-based rather than outage-driven. You will document all actions meticulously to transform findings into repeatable processes—and ultimately, into automated solutions—while refining deployment workflows to achieve seamless, predictable releases. Additionally, you will design, construct, and maintain scalable infrastructure capable of supporting hundreds of thousands of concurrent users, troubleshoot complex production issues across service layers, and strategically plan infrastructure expansion. Ideal candidates will demonstrate a cloud-first mindset, regardless of the underlying public cloud platform, and prioritize security in all initiatives. They will possess deep technical curiosity about systems—analyzing edge cases, failure modes, and implementation-specific behaviors—alongside strong proficiency in Linux and Windows environments. Familiarity with configuration management tools such as Ansible or Puppet, as well as programming expertise in languages like Python, Java, Go, or Node.js, is essential. The ability to work asynchronously, communicate clearly through documentation, and adopt a proactive “fix-it-now” attitude is critical.
Expérience
Experience with technologies such as Nginx, HAProxy, Docker, Kubernetes, Terraform, or equivalent platforms is highly valued. Potential projects include developing infrastructure automation with Ansible and Terraform, enhancing Prometheus monitoring or designing new metrics, assisting in software deployments, and leading the migration from traditional AWS virtual machines to cloud-native, containerized deployments on Kubernetes (EKS). You may also collaborate with product teams to define SRE key performance indicators, as this SRE function is in its early stages of development.
Formation / Diplômes
Secteur : Engineering and Information Technology
Your safety comes first — read before you go
Taf4All only lists job offers published by third parties. We are not the employer, we do not conduct these recruitments, and we cannot guarantee what happens once you make contact — you deal directly with the person or company behind the offer, at your own risk.
- •Never pay any amount of money — for a file, a training, a uniform, or an interview. A real employer never asks the candidate to pay.
- •Never send a photo of your ID card, passport, or banking details before you have physically verified the employer exists.
- •Always meet in a public place, during the day — never an isolated address, a private home, or a location you cannot verify in advance.
- •Tell a relative or friend exactly where you are going, with whom, and at what time — and share your live location if possible.
- •Search the company name online before going: a real business has a trace (website, reviews, other employees, an official address).
- •A salary that is far above the market rate for the position and the city is a red flag — be extra cautious.