Engineering Division SRE Platforms Software Engineering Vice President
Goldman SachsHyderabad, Telangana
it-jobs
Job Description
Description Site Reliability Engineer Vice President Site Reliability Engineering SRE is an engineering discipline that combines software and systems engineering to build and run scalable massively distributed faulttolerant systems At Goldman Sachs SRE is responsible for improving the availability and reliability of the firms most critical platform services and ensures they meet the requirements of our internal and external users It is also responsible for the firmwide policies and standards focused on firms digital resilience We are looking for engineers who are motivated to collaborate with our businesses to build and run sustainable production systems which can evolve and adapt to changes in our fastpaced global business environment The SRE team develops and maintains platforms and tools which help other Engineering teams in Goldman Sachs to build and operate reliable and resilient systems These systems span onpremises datacenters and multiple public cloud environments The platforms we offer include central logging monitoring agents and alerting and we provide tools to drive adoption and improvements to capacity planning operational readiness assessments production incident postmortems SLIs SLOs and deployment automation including canary releases The products and services we provide to our internal customers are used by thousands of engineers every day We believe that reliability is the most important feature of any system and we are devoted to giving our engineers the platforms and tools they need to build and operate reliable products Role Overview As a Site Reliability Engineer SRE at Goldman Sachs you will be a pivotal leader in ensuring the availability reliability and scalability of the firms most critical platform applications and services You will combine deep software and systems engineering expertise to architect build and run largescale massively distributed faulttolerant systems This role involves providing technical leadership mentoring senior engineers and collaborating closely with internal teams and executive stakeholders to build and operate sustainable production systems that can adapt to our dynamic global business environment You will drive a culture of continuous improvement championing the adoption of advanced SRE principles and best practices across the organization Responsibilities Strategic Reliability Performance Drive the strategic direction for availability scalability and performance of missioncritical applications and platform services ensuring alignment with firmwide objectives Architectural Leadership Lead the design build and implementation of highly available resilient and scalable infrastructure and application architectures Advanced Automation Tooling Architect and develop sophisticated platforms tools and automation solutions to eliminate toil optimize operational workflows and enhance deployment processes across the enterprise Complex Incident Management PostMortem Analysis Lead critical incident response conduct indepth root cause analysis for systemic issues and implement longterm preventative measures to significantly enhance system stability and resilience System Design Capacity Planning Partner with development teams to embed reliability into application design from inception provide expert system design consulting and lead comprehensive capacity planning initiatives for future growth Observability Insights Define and implement advanced monitoring high volume logging with multiuser query capabilities and tracing strategies to provide deep actionable insights into application performance infrastructure health and user experience Technical Vision Mentorship Provide technical vision lead complex technical projects conduct rigorous code reviews enforce SDLC best practices and actively mentor and develop senior and stafflevel engineers Technology Evaluation Adoption Stay at the forefront of industry trends and advancements evaluating and integrating cuttingedge tools and frameworks to significantly improve operational efficiency and reliability OnCall Leadership Participate in and lead oncall rotations providing expert guidance and handson support for critical system incidents Qualifications Experience Minimum of 6 years of handson experience in Site Reliability Engineering with a proven track record in architecting designing building and maintaining highly available scalable and faulttole
Get AI-Matched to This Job
Upload your resume and our AI will score how well you match this and thousands of similar roles.