Apply now »

Sr. SRE Specialist

Date:  23 Jul 2026
Company:  QualityAI
Country/Region:  IN
SRE Lead, Platform Reliability and OperationsNewLeaf Azure application platformThis is the role specification for a senior reliability leader who can combine SRE discipline, platform operations, and automation-first thinking across the NewLeaf Azure application estate.Role Summary• We are hiring a hands-on SRE Lead to own reliability across the full NewLeaf application platform on Azure.• This role covers customer-facing websites, backend services, APIs, integrations, event-driven workflows, observability, incident management, self-healing automation, on-call operations, and operational reporting.• The goal is to build a resilient operating model for the entire product ecosystem, not just the underlying cloud infrastructure.What "Platform" Means Here• Public and internal websites• Backend services and APIs• Azure-hosted application workloads• Integrations and event processing• Monitoring, alerting, dashboards, and reporting• Incident response, postmortems, and reliability improvements• Runbooks, operational standards, and self-healing automation• On-call rotation design and engineering readinessPrimary Responsibilities• Define the reliability strategy for the NewLeaf platform and keep it aligned with business priorities.• Own incident detection, triage, escalation, communication, and post-incident follow-through.• Build operational dashboards that show service health, user impact, and reliability trends.• Design self-healing and auto-remediation workflows for recurring failure modes.• Create and maintain runbooks, playbooks, and a living reliability rulebook.• Design and support a sustainable on-call rotation model for engineering teams.• Drive root-cause analysis and convert recurring issues into permanent fixes or automation.• Partner with application, integration, and platform teams to reduce operational risk across boundaries.Required Skill Set• Strong SRE or platform engineering background with direct ownership of production systems.• Deep incident management experience, including severity classification, incident command, and postmortem rigor.• Practical observability expertise across logs, metrics, traces, dashboards, alerting, and service health reporting.• Ability to build or influence automation for detection, triage, remediation, and follow-up tasks.• Experience shaping on-call practices, escalation paths, and team readiness.• Comfort operating across application services, integrations, asynchronous processing, and release workflows.• Clear communication skills for translating technical risk into operational and executive language.• Ability to coach engineers and raise the maturity of a team without creating unnecessary process overhead.Azure and Platform Stack Familiarity• Microsoft Azure operations and governance.• Azure Monitor, Application Insights, and Log Analytics.• Azure Service Bus, Event Grid, Event Hub, and timer-driven/background processing patterns.• Azure Functions, App Service, and other cloud-hosted application runtimes.• Key Vault, Storage, and identity-aware service configuration.• CI/CD pipelines and safe deployment practices, including release validation and rollback readiness.• Infrastructure-as-code tooling such as Bicep or Terraform, where used by the team.• Kusto Query Language and operational reporting from telemetry data.Agentic Operations and Automation• Use agents or automation to assist with analysis, incident summarization, runbook execution, and follow-up work.• Treat automation as a reliability control, not a novelty layer.• Apply safety checks, approvals, logging, and rollback paths before any automated remediation is allowed to act on production systems.• Continuously improve detection and remediation based on real incident patterns and post-incident learning.Leadership and Collaboration• Partner with engineering leaders to improve reliability without slowing delivery unnecessarily.• Lead cross-team incident coordination when production issues span multiple systems.• Establish practical standards for uptime, alert quality, runbook quality, and operational readiness.• Help build a culture where reliability is owned by every team, not delegated to a single hero.Nice-to-Have Experience• Experience introducing self-healing or auto-remediation in production.• Experience with executive dashboards and operational scorecards.• Experience in multi-team or multi-product environments with complex dependencies.• Background in Azure-native environments, distributed systems, and event-driven architectures.• Familiarity with AI-assisted operations or agent-based workflows in a safe enterprise setting.What Success Looks Like• Incidents are detected faster and resolved faster.• Recurring reliability issues are systematically eliminated or automated away.• On-call is sustainable, documented, and well supported.• Engineers have clear runbooks and confidence during outages.• Leadership has accurate visibility into platform health and risk.• Reliability becomes a repeatable operating discipline instead of a hero exercise.

Apply now »