ON

Site Reliability Engineering Specialist

Ontrac Solutions
📍 AntananarivoCDI🗓️ about 1 month ago

Job Description

Antananarivo
CDI

Your safety comes first — read before you go. Taf4All only lists job offers published by third parties. We are not the employer, we do not conduct these recruitments, and we cannot guarantee what happens once you make contact — you deal directly with the person or company behind the offer, at your own risk.

Our client’s Cloud Operations team is growing its Site Reliability Engineering (SRE) function to enhance the reliability and performance of user-facing services and production systems. This role demands a unique combination of hands-on operational expertise and software development proficiency, where engineers apply robust engineering principles, operational rigor, and automation to both infrastructure and code. The team focuses on core systems domains—including networking, Linux kernel optimization, scaling strategies, algorithmic efficiency, and distributed system architectures. In this position, you will participate in an on-call rotation to address production availability incidents, collaborate with service engineers to resolve customer-facing issues, and proactively leverage on-call time to mitigate potential incidents before they occur. Responsibilities include managing infrastructure through tools like Ansible, Puppet, Terraform, and Kubernetes, ensuring monitoring and alerting systems are symptom-based rather than outage-driven. You will document all actions meticulously to transform findings into repeatable processes—and ultimately, into automated solutions—while refining deployment workflows to achieve seamless, predictable releases. Additionally, you will design, construct, and maintain scalable infrastructure capable of supporting hundreds of thousands of concurrent users, troubleshoot complex production issues across service layers, and strategically plan infrastructure expansion. Ideal candidates will demonstrate a cloud-first mindset, regardless of the underlying public cloud platform, and prioritize security in all initiatives. They will possess deep technical curiosity about systems—analyzing edge cases, failure modes, and implementation-specific behaviors—alongside strong proficiency in Linux and Windows environments. Familiarity with configuration management tools such as Ansible or Puppet, as well as programming expertise in languages like Python, Java, Go, or Node.js, is essential. The ability to work asynchronously, communicate clearly through documentation, and adopt a proactive “fix-it-now” attitude is critical.

Expérience

Experience with technologies such as Nginx, HAProxy, Docker, Kubernetes, Terraform, or equivalent platforms is highly valued. Potential projects include developing infrastructure automation with Ansible and Terraform, enhancing Prometheus monitoring or designing new metrics, assisting in software deployments, and leading the migration from traditional AWS virtual machines to cloud-native, containerized deployments on Kubernetes (EKS). You may also collaborate with product teams to define SRE key performance indicators, as this SRE function is in its early stages of development.

Formation / Diplômes

Secteur : Engineering and Information Technology

---

**

[Click the Apply button below to apply, and Create my CV to build a CV tailored to this offer, professionally]

Ready to apply?

Does my CV fit this offer? Free diagnosis

Get seen by recruiters

Publish your CV on the Candidates page — the place recruiters browse directly to find profiles like yours.

See the Candidates page

You are recruiting?

Find candidates on TAF4ALL

🚀 Boost your application

Stand out with a professional CV and a personalized cover letter generated by AI in 3 minutes. 3 Ingénieur position(s) posted this week in Madagascar.

🇨🇦

Canadian employers are also recruiting in Africa

Real offers from Canadian employers who are explicitly looking for candidates outside Canada. No agency, no fees — you apply yourself.

Expert Application Advice

Our client’s Cloud Operations team is growing its Site Reliability Engineering (SRE) function to enhance the reliability and performance of user-facing services and production systems. This role demands a unique combination of hands-on operational expertise and software development proficiency, where engineers apply robust engineering principles, operational rigor, and automation to both infrastructure and code. The team focuses on core systems domains—including networking, Linux kernel optimization, scaling strategies, algorithmic efficiency, and distributed system architectures. In this position, you will participate in an on-call rotation to address production availability incidents, collaborate with service engineers to resolve customer-facing issues, and proactively leverage on-call time to mitigate potential incidents before they occur. Responsibilities include managing infrastructure through tools like Ansible, Puppet, Terraform, and Kubernetes, ensuring monitoring and alerting systems are symptom-based rather than outage-driven. You will document all actions meticulously to transform findings into repeatable processes—and ultimately, into automated solutions—while refining deployment workflows to achieve seamless, predictable releases. Additionally, you will design, construct, and maintain scalable infrastructure capable of supporting hundreds of thousands of concurrent users, troubleshoot complex production issues across service layers, and strategically plan infrastructure expansion. Ideal candidates will demonstrate a cloud-first mindset, regardless of the underlying public cloud platform, and prioritize security in all initiatives. They will possess deep technical curiosity about systems—analyzing edge cases, failure modes, and implementation-specific behaviors—alongside strong proficiency in Linux and Windows environments. Familiarity with configuration management tools such as Ansible or Puppet, as well as programming expertise in languages like Python, Java, Go, or Node.js, is essential. The ability to work asynchronously, communicate clearly through documentation, and adopt a proactive “fix-it-now” attitude is critical.

Expérience

Experience with technologies such as Nginx, HAProxy, Docker, Kubernetes, Terraform, or equivalent platforms is highly valued. Potential projects include developing infrastructure automation with Ansible and Terraform, enhancing Prometheus monitoring or designing new metrics, assisting in software deployments, and leading the migration from traditional AWS virtual machines to cloud-native, containerized deployments on Kubernetes (EKS). You may also collaborate with product teams to define SRE key performance indicators, as this SRE function is in its early stages of development.

Formation / Diplômes

Secteur : Engineering and Information Technology

Career advice powered by Taf4All

Your safety comes first — read before you go

Taf4All only lists job offers published by third parties. We are not the employer, we do not conduct these recruitments, and we cannot guarantee what happens once you make contact — you deal directly with the person or company behind the offer, at your own risk.

  • Never pay any amount of money — for a file, a training, a uniform, or an interview. A real employer never asks the candidate to pay.
  • Never send a photo of your ID card, passport, or banking details before you have physically verified the employer exists.
  • Always meet in a public place, during the day — never an isolated address, a private home, or a location you cannot verify in advance.
  • Tell a relative or friend exactly where you are going, with whom, and at what time — and share your live location if possible.
  • Search the company name online before going: a real business has a trace (website, reviews, other employees, an official address).
  • A salary that is far above the market rate for the position and the city is a red flag — be extra cautious.

You might also be interested in

ON

Reliability Engineer for Infrastructure Operations

Ontrac Solutions·Antananarivo

Our client’s Cloud Operations team is growing its Site Reliability Engineering (SRE) function to enhance the stability and efficiency of user facing services and production systems. This role demands a balanced combination of hands on operational expertise and software development proficiency, where engineers apply rob

CDIabout 1 month ago
JO

Anti-Money Laundering Specialist

📊 Autres informations About us: Paysera is the first fintech company in Lithuania and an EU licensed e money institution. We provide fast, convenient, and affordable financial services globally. Our services range from a payment gateway for e shops, a finance management app, money transfers, and a parcel locker networ

CDI6 days ago
EN

Human Resources Consulting Specialist

📋 Missions principales Missions principals : Piloted LE processes de recruitment pour vote Porterville clients, en analysand l’s begins, en identified et evacuant l’s profile, et en accompanist l’s (suggestion limit reached) (suggestion limit reached)’à (suggestion limit reached) (suggestion limit reached). Assurer LE

CDIabout 1 month ago
YA

Specialist : Survey and Upselling

NOUVEAU
YAS·Antananarivo

Description non disponible. Consultez l offre sur le site d origine pour plus de détails

CDIabout 16 hours ago