QA Lead - Cloud/AI/Infrastructure
NeysaMumbai, Maharashtra
it-jobs
Job Description
About the Role Neysa is building Velocis — an AI Infrastructure and PaaS platform that powers inference, orchestration, and multi-tenant GPU workloads at scale. This is not a conventional software QA role. The QA Lead will own the end-to-end quality strategy for Velocis — a distributed, Kubernetes-native platform spanning IaaS provisioning, AI-PaaS services, multi-tenant control planes, and real-time inference APIs. This is a hands-on leadership role: you will set the quality bar, build the team, define the toolchain, and be accountable for what ships. Key Responsibilities Quality Strategy & Engineering Leadership - Define and execute Neysa's Quality Engineering strategy across the Velocis platform - Establish QA processes, standards, and release governance across engineering teams - Build and mentor a high-performing QA organization; establish the hiring bar and career ladder - Drive a quality-first culture across Engineering, Product, and Platform Operations - Partner with Engineering Managers, Product, and SRE on release readiness and go/no-go decisions Test Planning & Execution - Own end-to-end test strategy for all Velocis platform releases - Review PRDs, technical designs, and architecture documents for testability — early, not late - Define test coverage across: Functional, Integration, Regression, Performance, Load, Security, API, and End-to-End testing - Apply risk-based testing to prioritize coverage on high-blast-radius surfaces Automation Excellence - Define automation roadmap, framework architecture, and coverage targets - Build scalable automation frameworks integrated into CI/CD pipelines - Drive regression automation coverage to measurable targets — own the metric, not just the plan - Instrument automation effectiveness: flakiness rate, coverage delta per release, mean time to detect Platform & Infrastructure Testing Lead quality across Velocis-specific surfaces: - Kubernetes-based PaaS and control plane workflows - AI/ML inference services (LLM APIs, model serving endpoints, non-deterministic output validation) - Multi-tenant isolation, RBAC, and provisioning correctness - Infrastructure provisioning workflows (IaaS, CloudStack abstraction, GPU allocation) - API gateway, microservices, and inter-service contract testing - User portals, dashboards, and operator control planes - Chaos engineering and fault injection — validate graceful degradation, not just happy paths - Observability-driven QA: use platform telemetry (metrics, traces, logs) as a first-class testing signal Quality Metrics & Governance Own and report quality KPIs to engineering leadership: - Defect Leakage Rate and Escaped Defects per release - Regression Defect Density - Automation Coverage % (with trend, not just snapshot) - Release Readiness Score - P95 API Latency Regression Detection - AI Inference Correctness Drift (output quality trend across model versions) - MTTR for Production Issues - Customer-Reported Defect Rate - Platform Reliability and SLA Adherence Release Management & Governance - Define and enforce release quality gates ,hard stops, not suggestions - Lead defect triage, root cause reviews, and post-mortems - Participate in Go/No-Go decisions with data-backed recommendations Customer & Production Quality - Analyze production incidents, identify systemic quality gaps, and drive preventive action - Partner with Support and SRE on issue resolution and observability tooling - Proactively surface quality risks before customers do AI-Augmented QA (What sets this role apart) Neysa expects this leader to actively drive AI-native quality practices — not as a future roadmap item, but as part of how the team operates now: - Use LLMs to generate test cases from PRDs, API specs, and architecture docs - Build AI-assisted test coverage gap analysis into the QA workflow - Apply AI-based log anomaly detection and failure clustering to accelerate RCA - Develop frameworks for testing non-deterministic AI outputs — semantic correctness, regression across model versions, and prompt-response consistency - Evaluate and adopt AI QA tooling (test generation, visual regression, self-healing locators) where they reduce manual overhead without sacrificing reliability - Contribute to Neysa's internal thinking on what "quality" means for agentic AI workloads Required Qualifications Experience - 10–14 years of Software QA experience, with depth in distributed or infrastructure products - Minimum 3–4 years leading QA teams (hiring, mentoring, performance management) - Experience in SaaS, Cloud Infrastructure, Platform Engineering, or Enterprise Software - Hands-on in Agile/Scrum environments; comfortable with fast release cadences Technical Skills — Required - API Testing: Postman, RestAssured, Swagger/OpenAPI contract testing - UI Automation: Playwright (preferred), Selenium, or Cypress - Performance & Load Testing: k6, JMeter, or Locust - Test Automation Framework Design (Python or Java-based) - CI/CD Integration: GitHub Actions, Jenkins, or equivalent - Kubernetes and containerized workloads — not just awareness, hands-on - Microservices and distributed systems testing - SQL and database correctness testing - Git-based workflows and test-as-code practices Technical Skills — Strong Preference - AI/ML platform or LLM API testing experience - GPU infrastructure or ML serving layer exposure - Chaos engineering tools (LitmusChaos, Gremlin, or equivalent) - Infrastructure as Code validation (Terraform, Ansible) - Security testing fundamentals (OWASP, API fuzzing, pen-test concepts) - Observability tooling: Grafana, Prometheus, distributed tracing Leadership Skills - Team building, hiring bar-setting, and mentoring - Stakeholder management across Engineering, Product, and SRE - Executive-level quality reporting , can translate defect data into business risk language - Ability to influence engineering quality practices without direct authority - Release governance owns the process, not just the checklist Why This Role You'll be the QA lead at a company building infrastructure for the AI era , the kind of platform where a test missed in staging can mean GPU time burned, SLA breaches, or a broken tenant boundary for a paying customer. The work is technical, high-stakes, and genuinely novel. You won't be maintaining a legacy test suite. You'll be building quality engineering from the ground up, on a platform that doesn't have many precedents. Autofill from resume Save time by uploading your resume. (Only PDF or DOCX format supported) Loading...
Get AI-Matched to This Job
Upload your resume and our AI will score how well you match this and thousands of similar roles.