Site Reliability Engineer (SRE)

EmbarkingOnVoyage Digital Solutions Pvt LtdPune, Maharashtra₹700,000 – ₹1,800,000
Adzuna INPosted 2h agoOriginal Listing
it-jobs

Job Description

Job Title: Site Reliability Engineer (SRE) – AI & Cloud Infrastructure Location: Pune (Work From Office) Experience: 5–8 Years Employment Type: Full-Time About the Role We are looking for an experienced Site Reliability Engineer (SRE) to build and scale AI-powered reliability capabilities from the ground up. In this role, you will drive modern observability, automation, and cloud reliability initiatives while leveraging AI/ML for incident management, forecasting, and infrastructure optimization. You will own the end-to-end reliability strategy across cloud-native AWS environments, enabling high availability, performance, and operational excellence through automation, intelligent monitoring, and proactive engineering. Key Responsibilities - Design, implement, and manage highly available, scalable, and secure cloud infrastructure on AWS. - Build and maintain an end-to-end observability platform using Open Telemetry, Grafana, Datadog, CloudWatch, and related tools. - Implement AIOps capabilities, including: - LLM-assisted incident triage - AI-powered root cause analysis - ML-driven forecasting and anomaly detection - Intelligent alert correlation and noise reduction - Lead production incident management, on-call response, postmortems, and Root Cause Analysis (RCA). - Automate operational workflows using Infrastructure as Code (Terraform/CloudFormation) and CI/CD pipelines. - Drive infrastructure rightsizing, capacity planning, utilization analysis, and cloud cost optimization. - Build dashboards, SLOs, SLIs, and error budgets to improve service reliability. - Develop automation scripts using Python, Bash, or Go to eliminate manual operational tasks. - Monitor application and infrastructure health while ensuring high uptime and service performance. - Collaborate with Development, DevOps, Security, Platform Engineering, and Product teams across multiple time zones. - Establish operational best practices for monitoring, incident response, disaster recovery, and resilience engineering. - Maintain Linux-based production systems and troubleshoot OS, networking, storage, and performance issues. Required Skills & Qualifications - 5–8 years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure Engineering. - Strong experience with AWS services including EC2, ECS/EKS, Lambda, VPC, IAM, CloudWatch, RDS, Route 53, S3, and Auto Scaling. - Hands-on experience with Open Telemetry, Grafana, Datadog, Prometheus, or similar monitoring platforms. - Strong knowledge of Linux administration, networking, system performance tuning, and troubleshooting. - Experience with Infrastructure as Code using Terraform or CloudFormation. - Proficiency in scripting using Python, Bash, or Go. - Experience with Kubernetes and containerized workloads. - Strong understanding of CI/CD pipelines (GitHub Actions, Jenkins, GitLab CI, etc.). - Experience leading incident management, production support, and RCA processes. - Knowledge of SRE principles including SLIs, SLOs, and Error Budgets. - Experience implementing monitoring, logging, alerting, and observability frameworks. - Strong analytical, troubleshooting, and communication skills. Preferred Qualifications - Experience building or implementing AIOps solutions. - Exposure to Large Language Models (LLMs) for operational automation. - Experience with machine learning-based forecasting or anomaly detection. - Hands-on experience administering Adobe Experience Manager (AEM). - Experience managing Cloudflare CDN, WAF, DNS, and caching strategies. - Knowledge of FinOps, cloud cost optimization, and capacity planning. - AWS Solutions Architect, DevOps Engineer, or Kubernetes certifications are a plus. Skills:- AIOps, Large Language Models (LLM), Machine Learning (ML), Amazon Web Services (AWS), Cloud-Native Infrastructure, OpenTelemetry, grafana, Datadog, Responsive Design, Root Cause Analysis (RCA), Infrastructure Design & Management, Infrastructure Rightsizing, Capacity Planning, Cost Anomaly Detection, Linux administration, Systems Administration, Adobe Experience Manager (AEM) Administration, Cloudflare CDN and Cross-functional Collaboration

Get AI-Matched to This Job

Upload your resume and our AI will score how well you match this and thousands of similar roles.