● Own the performance analytics for our AI agent fleet — tracking deflection rates, autonomous closure rates, escalation triggers, confidence score distributions, and CSAT by resolution type to give us a precise, real-time picture of where agents are succeeding and where they’re not.
● Dig into cases that escape AI resolution — identifying the intent clusters, knowledge gaps, and routing failures that cause unnecessary escalations, and translating those findings into specific improvements for agent training and prompt refinement.
● Define and govern the KPI framework for AI productivity — including human-time-saved per resolved case, agent-assisted vs. fully autonomous closure ratios, and the quality metrics that validate whether AI resolution is genuinely satisfying customers or just closing tickets.
● Build and maintain real-time dashboards that surface AI agent health, case closure velocity, and deflection trends — giving the support leadership a live signal on how the AI operation is performing at any given moment.
● Run structured QA and eval cycles on AI agent outputs — reviewing response accuracy, tone calibration, and resolution completeness to catch quality drift before it shows up in CSAT or re-open rates.
● Prepare concise, evidence-based briefings for leadership on AI agent performance and productivity gains — translating what the data shows into clear narratives about where to invest next to push resolution rates higher
WHAT YOU BRING
● 3–6 years of experience in analytics or business intelligence, with at least some of that time spent analyzing AI, ML, or LLM powered systems in a production environment you’ve measured how models behave in the real world, not just in test sets.
● Hands on experience with AI agent evaluation you know how to design rubrics, run QA cycles, and define what “good resolution” actually means in a support context. You can calibrate quality, not just count tickets.
● Sharp analytical instincts you can tell the difference between a deflection rate that’s improving because agents are getting better and one that’s improving because frustrated users are giving up. You ask the next question, not just the first one.
● Clear communication you can turn a complex AI performance analysis into a crisp recommendation that a support leader can act on immediately, without needing a technical deep-dive to understand the point.
● Familiarity with prompt engineering or context engineering enough to have informed opinion on why an agent responded the way it did and what a better prompt structure would look like
Stop applying to ghosts.
OyaPilot surfaces only verified, real jobs, scores your fit, and tailors your application truthfully.