SRE and Operations

Company
Description
Job Summary: We are seeking a passionate Operations and Site Reliability Engineer to lead the evolution of service operations and the reliability transformation of our Sales Platform, ensuring 24x7 availability. Key Highlights: 1. Leadership in transformation and ambassadorship of SRE culture. 2. Management of hybrid environments and root cause analysis of operational failures. 3. Focus on operational excellence and incident reduction. DESCRIPTION **Who Are We?** We are **Kairós**, a different kind of technology company. Yes, we help our clients tackle digital and methodological transformation challenges—but what truly defines us is **how we do it**: by putting people at the center. We have been recognized as a **Great Place To Work® company** since 2019—for the sixth consecutive year—and ranked **5th in 2024**. Our values of **courage, empathy, innovation, and joy** guide how we work every day. We are over 850 people (developers, data engineers, strategists…) and now looking for someone to join this brilliant and diverse team. At Kairós, you’re not just another number. We know you by name, care about your personal balance, and want you to grow with us. **Your Mission at Kairós** We seek a passionate **Operations and Site Reliability Engineer** to lead the evolution of service operations and the reliability transformation of our Sales Platform. Working side-by-side with Development and Service teams, your goal will be to guarantee absolute availability of a 24x7 platform serving 17 countries—achieving a true shift from a reactive to a proactive model. * **Leadership in Transformation:** You will actively participate in executing and ensuring the success of the service transformation plan. * **SRE Culture Ambassador:** You will introduce a culture of continuous improvement across development teams and act as a service ambassador to ensure best practices under the motto: *Everyone is service!* * **Hybrid Environment Management:** You will manage production and pre\-production environments based on hybrid architectures (On\-Premise and Cloud). * **Problem Analysis and Resolution:** You will conduct root cause analysis (RCA) of operational failures and deliver definitive solutions to prevent recurrence, leading *Problem Management*. * **Operational Excellence:** You will ensure compliance with SLAs, metrics, and KPIs—defining functional monitoring and escalation paths to improve incident detection. * **Support Optimization:** You will serve as second-level support, managing escalated incidents and developing procedures to empower first-level support to resolve issues autonomously. * **Incident Reduction:** You will continuously strive to reduce incidents as close to zero as possible—aligning service with the Delivery Life Cycle. * **Availability:** You will participate in compensated 24x7 on-call support shifts to ensure platform resilience. **What Will Make You Successful** #### **Experience and Competencies** * **Solid track record:** At least **4 years of IT experience**, preferably in corporate environments, as an IT Operations Support Analyst. * **Improvement mindset:** Customer-oriented approach, proactivity, autonomy, and a continuous improvement mindset. * **Teamwork:** Ability and initiative to collaborate effectively within multicultural teams accustomed to constant change. * **Adaptability:** Motivation to learn new methodologies and clear results orientation. * English **B2** #### **Technical Proficiency (Stack)** * **Cloud & Infrastructure:** Solid experience with **Azure**. Proficiency in virtualization and Infrastructure-as-Code (IaC), using tools such as **Docker** and **Kubernetes**. * **Systems:** Administration of servers and **Linux** operating systems. * **Automation:** Scripting skills using **Bash, Python, or Java**. * **Monitoring and Observability:** Experience with logging and monitoring tools such as **Prometheus, Grafana, ELK Stack (Elasticsearch, Logstash, Kibana)**, among others. * **Databases (Intermediate Level):** Knowledge of **SQL** and **MongoDB**. * **Development Skills (Intermediate Level):** Understanding and foundational knowledge of **Java SpringBoot, Angular, Ionic, Firebase**, and Web Services (**SOAP, REST**). * **Tool Management:** Experience using **Remedy, JIRA Service Desk, Jira**, and **Confluence**. * **Networking:** Solid networking knowledge. **Benefits That Make a Difference** **Meaningful Professional Growth** * Ongoing corporate training. * Free **English classes** in small groups. * Continuous mentoring and personalized guidance. * Participation in conferences, tech events, and communities. **Real Balance and Flexibility** * **23 vacation days** \+ December 24 and 31 off. * 1 day for professional events and **3 flexible conciliation days**. * **100% remote work:** You’ll work remotely with some flexibility to adapt your schedule—though this may vary depending on your project. **Personal Benefits and Care** * **Flexible compensation** (Cobee) for meals, transport, training, and childcare. * **Health insurance** with Cigna, Adeslas, or Sanitas. * **Wellness program** supporting emotional and psychological well-being. **\#TRY \#THINK \#FEEL \#ENJOY** At Kairós, we firmly believe in the importance of equality and diversity in the workplace. We are committed to creating an inclusive and respectful environment where all employees feel valued and have the opportunity to reach their full potential—regardless of race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, or any other characteristic protected by law. **Apply now and join the Kairoseros and Kairoseras team!**
Posted by

David Muñoz
Indeed · HR