ON

Reliability Engineer for Infrastructure Operations

Ontrac Solutions
📍 AntananarivoCDI🗓️ about 1 month ago

Job Description

Antananarivo
CDI

Your safety comes first — read before you go. Taf4All only lists job offers published by third parties. We are not the employer, we do not conduct these recruitments, and we cannot guarantee what happens once you make contact — you deal directly with the person or company behind the offer, at your own risk.

Our client’s Cloud Operations team is growing its Site Reliability Engineering (SRE) function to enhance the stability and efficiency of user-facing services and production systems. This role demands a balanced combination of hands-on operational expertise and software development proficiency, where engineers apply robust engineering practices, operational rigor, and advanced automation to both infrastructure and code. The team focuses on core systems engineering, including networking, Linux kernel optimization, scalability solutions, algorithmic efficiency, and distributed systems’ architecture. In this position, you will engage in an on-call rotation to address production availability incidents, providing support to service engineers during customer-facing issues. Your on-call responsibilities will extend beyond incident response—you will proactively identify and mitigate potential problems to prevent disruptions. Utilizing tools like Ansible, Puppet, Terraform, and Kubernetes, you will manage infrastructure deployment and configuration. Monitoring and alerting systems will be designed to detect early symptoms of issues rather than waiting for outages to occur. Every action taken will be meticulously documented, ensuring that insights are converted into repeatable processes and ultimately into automated solutions. Your contributions will focus on refining deployment workflows to achieve seamless, predictable operations, designing and maintaining scalable infrastructure capable of supporting hundreds of thousands of concurrent users, and resolving complex production issues across multiple layers of the technology stack. Additionally, you will play a key role in planning and scaling infrastructure to meet evolving demands.

Ideal candidates will demonstrate a cloud-first mindset, regardless of the specific public cloud platform, and prioritize security in every aspect of their work. A deep understanding of systems—including edge cases, failure modes, and performance behaviors—is essential, along with proficiency in Linux and Windows environments. Experience with configuration management tools such as Ansible or Puppet, and strong programming skills in languages like Python, Java, Go, or Node.js, are required. You should thrive in asynchronous collaboration, ensuring clear documentation to prevent knowledge silos. A proactive, solution-oriented attitude is critical—when you encounter broken systems, you will take ownership and resolve them efficiently. Familiarity with technologies such as Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar platforms is highly desirable. Potential projects include automating infrastructure with Ansible and Terraform, enhancing Prometheus monitoring or expanding metric collection, assisting release teams in deploying and troubleshooting application software, and leading the migration from legacy AWS virtual machines to cloud-native, containerized deployments on Kubernetes (EKS). You may also collaborate

Formation / Diplômes

Secteur : Engineering and Information Technology

---

**

[Click the Apply button below to apply, and Create my CV to build a CV tailored to this offer, professionally]

Ready to apply?

Does my CV fit this offer? Free diagnosis

Get seen by recruiters

Publish your CV on the Candidates page — the place recruiters browse directly to find profiles like yours.

See the Candidates page

You are recruiting?

Find candidates on TAF4ALL

🚀 Boost your application

Stand out with a professional CV and a personalized cover letter generated by AI in 3 minutes. 3 Ingénieur position(s) posted this week in Madagascar.

🇨🇦

Canadian employers are also recruiting in Africa

Real offers from Canadian employers who are explicitly looking for candidates outside Canada. No agency, no fees — you apply yourself.

Expert Application Advice

Our client’s Cloud Operations team is growing its Site Reliability Engineering (SRE) function to enhance the stability and efficiency of user-facing services and production systems. This role demands a balanced combination of hands-on operational expertise and software development proficiency, where engineers apply robust engineering practices, operational rigor, and advanced automation to both infrastructure and code. The team focuses on core systems engineering, including networking, Linux kernel optimization, scalability solutions, algorithmic efficiency, and distributed systems’ architecture. In this position, you will engage in an on-call rotation to address production availability incidents, providing support to service engineers during customer-facing issues. Your on-call responsibilities will extend beyond incident response—you will proactively identify and mitigate potential problems to prevent disruptions. Utilizing tools like Ansible, Puppet, Terraform, and Kubernetes, you will manage infrastructure deployment and configuration. Monitoring and alerting systems will be designed to detect early symptoms of issues rather than waiting for outages to occur. Every action taken will be meticulously documented, ensuring that insights are converted into repeatable processes and ultimately into automated solutions. Your contributions will focus on refining deployment workflows to achieve seamless, predictable operations, designing and maintaining scalable infrastructure capable of supporting hundreds of thousands of concurrent users, and resolving complex production issues across multiple layers of the technology stack. Additionally, you will play a key role in planning and scaling infrastructure to meet evolving demands.

Compétences requises

Ideal candidates will demonstrate a cloud-first mindset, regardless of the specific public cloud platform, and prioritize security in every aspect of their work. A deep understanding of systems—including edge cases, failure modes, and performance behaviors—is essential, along with proficiency in Linux and Windows environments. Experience with configuration management tools such as Ansible or Puppet, and strong programming skills in languages like Python, Java, Go, or Node.js, are required. You should thrive in asynchronous collaboration, ensuring clear documentation to prevent knowledge silos. A proactive, solution-oriented attitude is critical—when you encounter broken systems, you will take ownership and resolve them efficiently. Familiarity with technologies such as Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar platforms is highly desirable. Potential projects include automating infrastructure with Ansible and Terraform, enhancing Prometheus monitoring or expanding metric collection, assisting release teams in deploying and troubleshooting application software, and leading the migration from legacy AWS virtual machines to cloud-native, containerized deployments on Kubernetes (EKS). You may also collaborate

Formation / Diplômes

Secteur : Engineering and Information Technology

Career advice powered by Taf4All

Your safety comes first — read before you go

Taf4All only lists job offers published by third parties. We are not the employer, we do not conduct these recruitments, and we cannot guarantee what happens once you make contact — you deal directly with the person or company behind the offer, at your own risk.

  • Never pay any amount of money — for a file, a training, a uniform, or an interview. A real employer never asks the candidate to pay.
  • Never send a photo of your ID card, passport, or banking details before you have physically verified the employer exists.
  • Always meet in a public place, during the day — never an isolated address, a private home, or a location you cannot verify in advance.
  • Tell a relative or friend exactly where you are going, with whom, and at what time — and share your live location if possible.
  • Search the company name online before going: a real business has a trace (website, reviews, other employees, an official address).
  • A salary that is far above the market rate for the position and the city is a red flag — be extra cautious.

You might also be interested in

PA

DEVOPS & INFRASTRUCTURE ENGINEER -réf:DEV_ENG_TANA

PAOSITRA Finances·Antananarivo

Le/la DevOps Infrastructure Engineer contribue à la conception, au déploiement, à l’administration, à la sécurisation et à l’exploitation des infrastructures supportant la plateforme interne / Apache Fineract. 🎯 Responsabilités Administrer les environnements Linux et les serveurs applicatifs. Participer à la mise en p

CDI5 days ago
ON

Site Reliability Engineering Specialist

Ontrac Solutions·Antananarivo

Our client’s Cloud Operations team is growing its Site Reliability Engineering (SRE) function to enhance the reliability and performance of user facing services and production systems. This role demands a unique combination of hands on operational expertise and software development proficiency, where engineers apply ro

CDIabout 1 month ago
OP

Operations Intern

Operation Smile·Antananarivo

Operation Smile Madagascar is currently accepting applications for internship positions in several key areas, including Finance Operations, Program/Project Management, and Monitoring Evaluation. These internships will be available for a duration of 3 to 6 months, commencing in September 2026. Candidates must have compl

CDI28 days ago
MA

RESPONSABLE DES OPERATIONS

📋 Missions principales Piloter et coordonner les missions et interventions du cabinet Planifier les ressources, échéances et livrables Assurer le suivi opérationnel et la qualité des prestations Être l’interlocuteur privilégié des clients et veiller à leur satisfaction Développer les opportunités commerciales et les v

CDI5 days ago