SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE)

Company
Description
Summary: Seeking an experienced Site Reliability Engineer to enhance enterprise cloud environments through automation, observability, and operational excellence, ensuring high availability and resilience. Highlights: 1. Build Highly Reliable Cloud Platforms at Enterprise Scale 2. Modern Site Reliability Engineering culture 3. High-impact role supporting mission-critical platforms #### **SITE RELIABILITY ENGINEER (SRE) (HYBRID / REMOTE BRAZIL)** Portuguese company hires for hybrid position Location: Brazil (any location) * ️ Only candidates already based in Brazil will be considered Work Model: Hybrid for candidates living in state capitals and Remote for candidates living in countryside/cities outside the state capitals ️ Language Requirements: Advanced/Fluent English – **Mandatory**Mandatory (there will be direct contact with an international client) Seniority: Senior (6\+ years) Compensation: Please inform your salary expectations when applying. * ️ Instructions: Please send your CV in English and make sure to include all skills and experience that match the requirements of the opportunity. This will significantly increase your chances of success. #### **Build Highly Reliable Cloud Platforms at Enterprise Scale** We are looking for an experienced Site Reliability Engineer (SRE) to improve the reliability, scalability, performance, and operational excellence of enterprise cloud environments. You will work closely with development, platform, and infrastructure teams to automate operations, improve observability, optimize deployments, and ensure high availability across mission\-critical systems. If you enjoy solving complex infrastructure challenges through engineering, automation, and modern cloud technologies, this opportunity is for you. #### **The Professional We Are Looking For** We are seeking a highly skilled Site Reliability Engineer with extensive experience in cloud infrastructure, DevOps practices, automation, and distributed systems. The ideal candidate combines strong infrastructure knowledge with software engineering principles, helping organizations improve operational efficiency through automation, Infrastructure as Code (IaC), observability, and continuous improvement. You should be comfortable working in high\-availability environments, responding to critical incidents, and continuously enhancing platform resilience. #### **Key Responsibilities** You will be responsible for: **Automation \& Operational Excellence** * Automate operational processes, deployments, and infrastructure provisioning. * Develop internal automation tools using **Python** and **Bash**. * Design, implement, and optimize CI/CD pipelines. * Reduce repetitive operational activities through automation. **Cloud Infrastructure \& Infrastructure as Code** * Provision and manage cloud infrastructure. * Implement Infrastructure as Code using **Terraform** (preferred) and other IaC tools. * Ensure scalability, resilience, and high availability. * Design auto\-scaling and self\-healing solutions. **Reliability \& Performance** * Monitor applications and infrastructure using metrics, logs, and distributed tracing. * Define and monitor **SLIs, SLOs, and SLAs**. * Perform capacity planning. * Troubleshoot performance bottlenecks and system issues. **Incident Management** * Respond to critical production incidents. * Conduct Root Cause Analysis (RCA). * Drive continuous improvement initiatives. **Observability** * Implement monitoring, alerting, and observability platforms. * Build operational dashboards. * Improve end\-to\-end system visibility. **Platform Engineering** * Collaborate closely with software engineering teams. * Support DevOps and SRE best practices. * Improve deployment processes, system architecture, and application resilience. **Business Continuity** * Support Disaster Recovery strategies. * Execute load and stress testing. * Implement Chaos Engineering practices. * Improve operational resilience. **Governance** * Define operational standards. * Document infrastructure and architecture. * Promote security best practices and DevSecOps principles. #### **Mandatory Requirements (Eliminatory)** * ️ **All requirements below are mandatory and eliminatory.** Candidates who cannot clearly demonstrate these qualifications in their CV are unlikely to proceed in the recruitment process. **Education** ✔ Bachelor's Degree in: * Computer Science * Engineering * or a related field #### **Mandatory Experience** * Advanced/Fluent English. * Proven experience as a **Site Reliability Engineer (SRE)**. * Strong background in: * Site Reliability Engineering * DevOps * Cloud Infrastructure * Cloud Engineering * Hands\-on experience with at least one major cloud platform: * AWS * Microsoft Azure * Google Cloud Platform (GCP) * Practical experience with: * Terraform (preferred) * Docker * Kubernetes * CI/CD Pipelines * Python or Bash * Linux/Unix * Strong experience with monitoring and observability platforms. * Incident Management and Production Support. * Strong troubleshooting and problem\-solving skills. * Experience building highly available and scalable systems. #### **Nice\-to\-Have Skills** The following qualifications will be considered a strong advantage: * Large\-scale enterprise environments. * Consulting experience. * Chaos Engineering. * Highly distributed architectures. * Advanced observability platforms. * Apache Kafka. * Financial Services or mission\-critical environments. * DevSecOps. * Platform Engineering. #### **What You'll Find in This Opportunity** * Enterprise\-scale cloud infrastructure projects. * Modern Site Reliability Engineering culture. * Cloud\-native and Platform Engineering initiatives. * Strong DevOps and Infrastructure as Code practices. * Collaboration with international engineering teams. * High\-impact role supporting mission\-critical platforms. * Modern engineering environment focused on automation and operational excellence. #### **Before Applying, Ask Yourself These 5 Questions** ✅ Have I worked as a **Site Reliability Engineer (SRE)** or in an equivalent role focused on **cloud reliability, automation, and operational excellence**? ✅ Does my CV clearly demonstrate hands\-on experience with **AWS, Azure, or GCP**, as well as **Terraform, Docker, Kubernetes, Linux, Python/Bash, and CI/CD pipelines**? ✅ Have I implemented **observability solutions**, managed production incidents, conducted **Root Cause Analysis (RCA)**, and worked with **SLIs, SLOs, and SLAs**? ✅ Do I have experience designing highly available, scalable, and resilient cloud environments using Infrastructure as Code and DevOps best practices? ✅ Am I fluent in English and comfortable collaborating with international engineering teams in mission\-critical production environments? If you answered **"No"** to one or more of these questions, we recommend carefully reviewing your fit before applying. * **️ Important** This position is intended for a **Senior Site Reliability Engineer (SRE)** with extensive experience in **cloud infrastructure, automation, observability, DevOps, and platform reliability**. Candidates whose experience is primarily focused on traditional infrastructure administration, system support, or operations without demonstrated expertise in **Infrastructure as Code, Kubernetes, cloud platforms, automation, CI/CD, and Site Reliability Engineering practices** are unlikely to meet the expectations for this role. #### **Keywords That Should Appear in Your CV** **Site Reliability Engineer, SRE, DevOps Engineer, Platform Engineer, Cloud Engineer, Infrastructure Engineer, Cloud Infrastructure, AWS, Amazon Web Services, Microsoft Azure, Google Cloud Platform, GCP, Terraform, Infrastructure as Code, IaC, Docker, Kubernetes, Linux, Unix, Python, Bash, Shell Scripting, CI/CD, Jenkins, GitLab CI, GitHub Actions, Monitoring, Observability, Prometheus, Grafana, ELK Stack, OpenTelemetry, Distributed Tracing, Logging, Metrics, SLI, SLO, SLA, Incident Management, Root Cause Analysis, RCA, Capacity Planning, Auto Scaling, Self\-Healing, Disaster Recovery, Chaos Engineering, DevSecOps, Platform Engineering, Apache Kafka, High Availability, Scalability, Enterprise Infrastructure, Financial Services** \#EP
Posted by

João Santos
Indeed · HR




