Job Description
Job Description We are seeking a highly motivated Senior DevOps Engineer with strong experience in designing, automating, and operating scalable, secure, and highly available cloud-native platforms. The role requires hands-on expertise across AWS, Kubernetes/EKS, Infrastructure as Code, CI/CD, observability, service mesh, databases, and production operations. The ideal candidate should also have a good understanding of Generative AI, LLM platforms, AI gateways, and practical AI adoption in DevOps/AIOps workflows. Roles & Responsibilities - Design, build, operate, and scale highly available production infrastructure on AWS, with a strong focus on reliability, security, automation, and performance. - Manage and optimize Kubernetes/EKS environments, including cluster upgrades, autoscaling, workload scheduling, networking, security, resource management, and high availability. - Build and maintain reusable Infrastructure as Code using Terraform and implement automated infrastructure provisioning across multiple environments. - Design and maintain robust CI/CD pipelines using GitLab CI/CD, Jenkins, or similar tools for automated build, testing, deployment, rollback, and release management. - Implement and operate containerized workloads using Docker, Kubernetes, and Helm, following deployment and configuration best practices. - Manage service-to-service communication and traffic using Istio, Kubernetes Gateway API, kgateway, Envoy-based gateways, or similar technologies. - Build and maintain centralized observability platforms using Prometheus, Grafana, Loki, Cortex, and related monitoring, logging, alerting, and tracing tools. - Troubleshoot complex production issues across Kubernetes, Linux, networking, DNS, HTTP/TLS, service mesh, load balancers, databases, and distributed microservices. - Support and operate data platforms such as MySQL, PostgreSQL, RDS, ScyllaDB, ClickHouse, and Redis, including backup, replication, monitoring, and performance troubleshooting. - Drive platform reliability and operational excellence through incident management, root cause analysis, automation, capacity planning, cost optimization, and continuous improvement. - Adopt AI and AIOps capabilities in DevOps workflows for troubleshooting, log analysis, alert correlation, automation, incident analysis, documentation, and engineering productivity. Required Skills - 5-6 years of hands-on experience in DevOps, Cloud Infrastructure, SRE, or Platform Engineering roles. - Strong knowledge of AWS services including EKS, EC2, VPC, ALB/NLB, CloudFront, Route 53, IAM, S3, RDS, KMS, Secrets Manager, and security services. - Strong experience with Kubernetes, Docker, Helm, Terraform, GitLab CI/CD, and Jenkins. - Hands-on knowledge of Istio, service mesh, API gateways, Kubernetes Gateway API, and Envoy-based architectures. - Strong experience with observability tools such as Prometheus, Grafana, Loki, Cortex, and similar monitoring/logging platforms. - Strong Linux administration and troubleshooting skills, with scripting experience using Python and Shell. - Experience with relational and distributed databases including MySQL, PostgreSQL, RDS, ScyllaDB, ClickHouse, Redis, or similar technologies. - Familiarity with build and software development tools such as Git, Maven, Gradle, artifact repositories, and release processes. - Good understanding of Generative AI and LLM fundamentals, including prompts, tokens, context windows, model APIs, inference, latency, rate limits, and AI workload cost. - Familiarity with AI platforms and technologies such as OpenAI, Anthropic Claude, Google Gemini, AWS Bedrock, AI-gateway, or similar AI/LLM gateway solutions.
Get AI-Matched to This Job
Upload your resume and our AI will score how well you match this and thousands of similar roles.