Cloud Site Reliability Engineer
ExperianHyderabad, Telangana
it-jobs
Job Description
Job Description Overview: We are recruiting an experienced Cloud Engineering Site Reliability Engineer (SRE) to join our Cloud Engineering team. The Cloud Engineering SRE plays a critical role in designing, improving cloud platforms and infrastructure services that support business-critical products and services. The SRE works with application teams, architecture teams, security teams, and platform engineers. The SRE ensures the reliability, scalability, availability, performance, and security of cloud-based systems throughout their lifecycle. As a primary member of the Cloud Engineering function, the SRE drives through automation, observability, resilience engineering, and continuous improvement practices. Summary: As part of the Cloud Engineering team, participate in the design, implementation, operation, and optimization of cloud platform. Develop expertise in cloud-native technologies, infrastructure automation, reliability engineering. The SRE will work along with engineering, architecture, and business teams to improve platform reliability, reduce operational toil, implement observability solutions, and ensure services meet agreed service levels and customer expectations. You will be reporting to a Director. Key Responsibilities Cloud Infrastructure Engineering · Develop secure, highly available cloud infrastructure platforms. · Implement Infrastructure as Code (IaC) using industry-standard tooling. · Support cloud infrastructure lifecycle management including maintenance, optimization, and retirement. · Collaborate with architecture and engineering teams to ensure cloud solutions align with enterprise standards. · Help with capacity planning, performance tuning, and infrastructure optimization. Automation & DevOps · Automate repetitive operational activities to reduce manual effort and improve reliability. · Build CI/CD pipelines and deployment processes. · Develop self-healing capabilities and automated remediation mechanisms. · Promote Infrastructure as Code, GitOps, and cloud-native engineering practices. · Improve deployment reliability through testing, validation, and release automation. Observability & Compliance · Analyze system performance and identify opportunities to improve efficiency and reliability. · Create dashboards and operational metrics to support service health monitoring. · Ensure cloud environments comply with security, regulatory, and governance requirements. · Participate in vulnerability remediation and risk management activities. · Identify operational, technical, and security risks and lead mitigation plans. Collaboration & Technical Leadership · Work with architecture, security, networking, and infrastructure build teams. · Provide technical guidance on cloud reliability and operational best practices. · Contribute to engineering standards, operational frameworks, and platform strategy. · Stay informed of latest cloud technologies and industry best practices. Technologies You have experience with several of the following technologies: · AWS / Azure · Kubernetes / Docker · Terraform · GitHub Actions / Jenkins · Prometheus / Grafana · Datadog / Splunk · OS – Linux / Windows Server · Scripting – Python / Bash / PowerShell · REST APIs · High-level Networking and Security Services Certifications: · AWS Certified DevOps Engineer · Microsoft Certified: Azure DevOps Engineer Expert · Certified Kubernetes Administrator (CKA) · ITIL Foundation · Terraform Associate
Get AI-Matched to This Job
Upload your resume and our AI will score how well you match this and thousands of similar roles.