Site Reliability & DevOps Engineer

Mitra InnovationIndia
Adzuna INPosted -40m agoOriginal Listing
it-jobs

Job Description

Location: Sri Lanka (On-site / Hybrid) or India Type: Full-time Experience Level: 3+ Years About the Role We are seeking a Site Reliability & DevOps Engineer to support the operational resilience, automation, and security of our cloud infrastructure and modern data platforms. In this role, you will be responsible for maintaining CI/CD pipelines, managing infrastructure-as-code (IaC), executing credential rotation routines, and assisting with observability and chaos engineering practices. You will work closely with Senior SRE leads, Software Engineers, and Data Platform teams (Snowflake, dbt Cloud, Fivetran) within financial domain environments to support secure releases, incident response, and high-performance infrastructure. Key Responsibilities1. Infrastructure as Code (IaC) & Cloud Provisioning - Infrastructure Maintenance: Maintain and enhance existing Terraform modules and Ansible playbooks used for provisioning AWS resources, GitHub configurations, and dbt Cloud environments. - Environment & Key Management: Execute automated environment setups, credential rotations, AWS IAM role updates, RSA key management, and 1Password Vault configurations. - Specialized Infrastructure Tooling: Utilize phData Provisioning tools and custom infrastructure provisioning frameworks following team guidelines. 2. CI/CD & Release Operations - Pipeline Execution & Maintenance: Build and maintain continuous integration and delivery pipelines using tools such as ArgoCD, Jenkins, and GitHub/GitLab Actions. - Containerization & Deployment: Deploy and manage Kubernetes workloads, Helm charts, and Dockerized microservices across development, staging, and production environments. - Automated Testing Integration: Run performance and API testing mechanism integrations (e.g., JMeter, Newman) within deployment pipelines to validate software updates. 3. Reliability, Observability & Incident Response - SRE Practices: Participate in SRE operational workflows, including tracking Service Level Indicators/Objectives (SLIs/SLOs), monitoring error budgets, and attending blameless postmortems. - Resilience Testing: Assist senior team members in executing chaos engineering scenarios and failure-injection tests to validate application recovery. - Monitoring & Alerting: Help maintain observability tools (Prometheus, Elasticsearch/Kibana (ELK), AWS CloudWatch) and respond to automated system alerts. - Runbooks & Troubleshooting: Follow operational runbooks, perform initial root-cause analysis (RCA) on bug reports, and coordinate fixes with development teams. 4. Security & Stakeholder Collaboration - Financial Domain Security: Adhere to cybersecurity policies, execute routine vulnerability fixes, and follow compliance standards for financial systems. - Team Collaboration: Communicate effectively with internal software teams and clients, reporting daily progress to technical leads. Required Qualifications & Technical Skills - Experience: 3+ years of experience in software engineering with a focus on DevOps, Cloud Engineering, or Site Reliability Engineering. - Cloud & Networking: Working knowledge of AWS core services (EC2, S3, IAM, VPC), DNS, load balancers (ELB/ALB), and basic networking protocols (TCP/IP, VPN). - Infrastructure as Code (IaC): Hands-on experience writing or modifying Terraform and Ansible playbooks. - Containers & Orchestration: Hands-on experience with Docker, Kubernetes, and Helm. - CI/CD Tools: Experience working with ArgoCD, Jenkins, or Git-native CI/CD platforms (GitHub/GitLab Actions). - Programming & Scripting: Proficiency in Python or Shell scripting, with exposure to Java or .NET environments. - Security & Vaulting: Experience managing AWS IAM permissions, RSA key pairs, and secrets management tools (e.g., 1Password, Vault). Preferred / Domain-Specific Skills - Data Ecosystem Exposure: Basic exposure to Snowflake, Fivetran, dbt Cloud, or distributed data tools (e.g., Kafka, Cassandra). - Resilience Testing: Familiarity with chaos engineering concepts or tools (e.g., Gremlin, Chaos Mesh). - Documentation: Proven ability to author clear technical runbooks and post-incident summary logs. Key Performance Indicators (KPIs) - Deployment Reliability: High success rate and quick turnaround on routine pipeline execution and deployment tasks. - Ticket Resolution Time: Prompt initial triage and resolution of infrastructure and build failure tickets. - Automation Adherence: Consistent use of IaC (Terraform/Ansible) for routine environment configuration changes. - SLI/SLA Uptime: SLA compliance across supported cloud environments and data pipelines.

Get AI-Matched to This Job

Upload your resume and our AI will score how well you match this and thousands of similar roles.