Our client’s Cloud Operations team is growing its Site Reliability Engineering (SRE) function to enhance the stability and efficiency of user-facing services and production systems. This role demands a balanced combination of hands-on operational expertise and software development proficiency, where engineers apply robust engineering practices, operational rigor, and advanced automation to both infrastructure and code. The team focuses on core systems engineering, including networking, Linux kernel optimization, scalability solutions, algorithmic efficiency, and distributed systems’ architecture.
In this position, you will engage in an on-call rotation to address production availability incidents, providing support to service engineers during customer-facing issues. Your on-call responsibilities will extend beyond incident response—you will proactively identify and mitigate potential problems to prevent disruptions. Utilizing tools like Ansible, Puppet, Terraform, and Kubernetes, you will manage infrastructure deployment and configuration. Monitoring and alerting systems will be designed to detect early symptoms of issues rather than waiting for outages to occur. Every action taken will be meticulously documented, ensuring that insights are converted into repeatable processes and ultimately into automated solutions.
Your contributions will focus on refining deployment workflows to achieve seamless, predictable operations, designing and maintaining scalable infrastructure capable of supporting hundreds of thousands of concurrent users, and resolving complex production issues across multiple layers of the technology stack. Additionally, you will play a key role in planning and scaling infrastructure to meet evolving demands.
Ideal candidates will demonstrate a cloud-first mindset, regardless of the specific public cloud platform, and prioritize security in every aspect of their work. A deep understanding of systems—including edge cases, failure modes, and performance behaviors—is essential, along with proficiency in Linux and Windows environments. Experience with configuration management tools such as Ansible or Puppet, and strong programming skills in languages like Python, Java, Go, or Node.js, are required. You should thrive in asynchronous collaboration, ensuring clear documentation to prevent knowledge silos. A proactive, solution-oriented attitude is critical—when you encounter broken systems, you will take ownership and resolve them efficiently.
Familiarity with technologies such as Nginx, HAProxy, Docker, Kubernetes, Terraform, or similar platforms is highly desirable. Potential projects include automating infrastructure with Ansible and Terraform, enhancing Prometheus monitoring or expanding metric collection, assisting release teams in deploying and troubleshooting application software, and leading the migration from legacy AWS virtual machines to cloud-native, containerized deployments on Kubernetes (EKS). You may also collaborate
Formation / Diplômes
Secteur : Engineering and Information Technology
---
**
[Click the Apply button below to apply, and Create my CV to build a CV tailored to this offer, professionally]