Job Description
Job Description We are seeking a talented and motivated SRE Engineer III to join our dynamic team. In this role, you will execute a range of site reliability activities, ensuring optimal service performance, reliability, and availability. You will collaborate with cross-functional engineering teams to develop scalable, fault-tolerant, and cost-effective cloud services. If you are passionate about site reliability engineering and ready to make a significant impact, we would love to hear from you! Key Responsibilities: - Implement automation tools, frameworks, and CI/CD pipelines, promoting best practices and code reusability. - Enhance site reliability through process automation, reducing mean time to detection, resolution, and repair. - Identify and manage risks through regular assessments and proactive mitigation strategies. - Develop and troubleshoot large-scale distributed systems in both on-prem and cloud environments. - Deliver infrastructure as code to improve service availability, scalability, latency, and efficiency. - Monitor support processing for early detection of issues and share knowledge on emerging site reliability trends. - Analyze data to identify improvement areas and optimize system performance through scale testing.
Get AI-Matched to This Job
Upload your resume and our AI will score how well you match this and thousands of similar roles.