Job Description
About This Role As Manager, Incident Management at KIBO, you lead the team that owns Sev1 response for our global commerce platform. You run a team of Incident Managers, you own the Sev1 process itself — the severity bar, the playbook, the deadlines, and the escalation path — and you command Sev1 bridges yourself. You are accountable for what comes out of every incident. RCAs go to the client complete and on time. Sev1 KPIs are monitored and reported, and you know which way each one is moving. And the pattern across incidents gets found rather than lost: when the same cause keeps returning, you bring it to the CTO, CPO, and CCO with a clear read on what has to change in the platform, in the product, or in how we operate. About KIBO KIBO is a composable digital commerce platform for B2C, D2C, and B2B organizations who want to simplify the complexity in their businesses and deliver modern customer experiences. KIBO is the only modular, modern commerce platform that supports experiences spanning B2B and B2C Commerce, Order Management, and Subscriptions. Companies like Ace Hardware, Zwilling, Jelly Belly, Nivel, and Honey Birdette trust KIBO to bring simplicity and sophistication to commerce operations and deliver experiences that drive value. KIBO's cutting-edge solution is MACH Alliance Certified and has been recognized by Forrester, Gartner, IDC, Internet Retailer, and TrustRadius. KIBO has been named a leader in The Forrester Wave™: Order Management Systems, Q1 2025 and in the IDC MarketScape report "Worldwide Enterprise Headless Digital Commerce Applications 2024 Vendor Assessment." By joining KIBO, you will be part of a team of Kibonauts all over the world in a remote-friendly environment. Whether your job is to build, sell, or support KIBO's commerce solutions, we tackle challenges together with the approach of trust, growth mindset, and customer obsession. What You'll Do - Take your own shifts in the Incident Manager rotation — carry the page, triage, command the bridge, and write the RCA, on the same terms as your team - Own the team's operating model — coverage, rotation cadence, and handoffs — so Sev1 command is always claimed on time and never depends on one person being reachable - Hire, train, and coach three Incident Managers; certify each one through shadowing before they command a bridge alone - Own the severity standard — the Sev1 criteria, the per-service baseline thresholds, and the client-contact decision test — and keep it applied consistently across every IM and every shift - Be the escalation point when an IM's call is challenged, including client-driven pressure to open a bridge that does not meet the criteria; where an override is warranted, make sure it comes with an RCA that tests whether it was justified - Own the First Call playbook and the triage runbooks, and keep the rule that they only grow from real incidents — one incident, one entry - Run problem management: turn recurring incidents and RCA findings into permanent fixes with named owners and committed dates, and chase them to closed rather than to filed - Own Sev1 KPI monitoring and reporting — volume, deadline adherence, and repeat-cause trends — with a clear read on what is improving and what is not - Identify the patterns across Sev1s and take them to the CTO, CPO, and CCO as specific recommendations: what to fix in the platform, what to change in the product, and what to change in how we operate - Tune the severity thresholds against each service's real numbers so the bar reflects actual traffic patterns by hour and weekday, including tighter thresholds for top-tier accounts and peak periods - Partner with DevOps and Engineering on alert quality — reduce the pages that carry no signal, and close the gaps where a real outage produces no page at all - Own the paging configuration: rotation order and cadence, acknowledgement windows, automatic escalation timers, and corresponding automation, so no incident depends on one person being awake - Keep the IM roster, pod-to-account mapping, and escalation contacts published and current What You'll Need - 10+ years in incident management, production support, or technical operations, including 2+ years leading a team or an on-call function - Direct experience as Incident Commander on major incidents — you have run the bridge, not just reported on it afterward - Willingness to carry a page. This role is on call and takes rotation shifts. - Enough technical depth to follow a live debug and push on it — logs, APIs, queues, databases, cloud infrastructure — with the discipline not to take the work over - A track record of holding a severity standard under pressure, including telling a large account no and making the decision hold - Experience converting RCAs into completed engineering work, with evidence that the same incident stopped recurring - Working knowledge of ITIL incident and problem management, applied in practice rather than recited - Hands-on with observability and paging tooling (PagerDuty, Opsgenie, Grafana, Splunk, Prometheus, or equivalent) - Familiarity with cloud platforms (AWS, GCP, or Azure) and containerized environments (Kubernetes/Docker) - Writing and presence that hold up in front of an enterprise client's executive team Bonus - Built or materially rebuilt an incident management practice, not just operated one - Commerce, retail, or payments platform experience — peak season, checkout, payments, order flow - Use of enterprise-approved AI/GenAI tooling to accelerate log analysis and RCA drafting, with judgment about data sensitivity and validating what the tool produces KIBO Perks - Flexible schedule and time away programs - Paid company holidays and global volunteer day - Generous health, wellness, and benefit programs, including 401(k) match and pet insurance - Opportunity for impact, rapid career growth, and intellectual stimulation - Passionate, high-achieving teammates excited to help you succeed and learn - Company events and other activities (Holiday parties, Meet-ups, Volunteering) At KIBO we celebrate and support all differences. KIBO is proud to be an equal opportunity workplace. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital, disability, and veteran status.
Get AI-Matched to This Job
Upload your resume and our AI will score how well you match this and thousands of similar roles.