How Hugo Strengthened Persona-Driven Conversational AI at Scale
A persona-led chatbot has a harder job than a general assistant. It has to be accurate, and it has to stay in character, and when those two pull in opposite directions something has to give. Deciding which should give, across millions of conversations, is a human judgment problem.
A technology organization running a creator platform for persona-led conversational agents engaged Hugo to evaluate and strengthen the conversational quality of those personas. Across eight workflows the team delivered ~1.16 million annotations, reached 97.76% QA against a 90% target and 92.5% accuracy against 81.1%, at an average handling time 27% faster than target.
Why It Was Hard
Two workflows that look identical and are not
The sharpest challenge was structural. Preference and Factuality share the same inputs, the same layout and nearly the same steps: user prompt, bio and style directives, social memory facts, message history, model responses. What differs is the lens. Preference asks which response is better. Factuality asks which is true. A rater running the wrong lens produces confident, well-formed, useless data.







Multi-step reasoning over full histories
Ranking a response meant reading the whole conversation, categorising user intent from it, and only then judging the candidates. Context could not be skipped.

Efficiency that moved for reasons outside the work
Seed data shortages, tooling timeouts, skipped complex jobs and network instability all pulled at occupancy and non-handling time.
Rater harmony
With this density of subjective decisions, agreement between raters, QAs and client auditors is not a reporting metric. It is the product.
How Hugo Approached It
Upfront preparation
Workflow-specific SOPs, scoring rubrics, calibration packs and simulation exercises, built around persona depth, tone, preference logic and factual accuracy. A 24-hour feedback and RCA loop closed client-identified issues before they could recur.
Team structure, organised against cognitive load
176 raters and 43 QPEs were organised into domain-aligned units so context mastery was preserved rather than diluted. Frequent queue switching was replaced with a main plus fallback queue structure, and every rater was cross-skilled on exactly one fallback: flexibility without losing specialisation.
Four dedicated QA coaches ran targeted calibration pods at a 1:15 ratio, clarifying ambiguous cases and deepening mastery in the workflows where misinterpretation was most likely. Small pods outperformed broad training calls consistently.
A triple-layered review loop
- Rater review. Each morning, raters worked through the prior day’s audits, identified patterns and shared what they found.
- QA review. QAs validated consistency and surfaced the nuanced cases that guidelines had not yet resolved.
- Client QA reconciliation. Flagged discrepancies were reviewed inside 24 hours, so recalibration happened while the context was still fresh.
Daily audit-of-audit reviews reinforced scoring discipline and held QA above 97% from July 2024 onward.
Continuous calibration
Daily RCA ran across quality, utilization, occupancy and handling time to catch trends early. Audit-of-audit sessions were used specifically to refine the subjective rubrics and settle the Preference-versus-Factuality boundary. A structured communication cadence with the client kept tool and instruction changes from sitting unactioned.
Efficiency and reliability
New queues were tested early to catch timer and tooling errors before they hit production, best practice for error-prone jobs was reinforced, and escalation pathways were shortened to cut downtime.
Results
- ~1.16 million annotations across eight workflows.
- 97.76% QA, exceeding target by 7.76 points.
- 92.5% accuracy, exceeding goal by 11.4 points.
- Average handling time 1,428 seconds against a 1,974-second target, 27% faster.
- 97.75% occupancy with non-handling time at 1.96% after optimisation.
- A 43-week sustained delivery cadence.
What We Learned
Similar-looking workflows need explicit separation
Where two workflows share structure but differ in lens, calibration and worked examples are the only reliable fix. Rubric text alone will not hold the line.
Small coaching pods beat broad training
Nuance transfers when feedback is contextual and tailored. It does not transfer well on a group call.
Stability lowers cognitive load
Stabilising queue assignments improved quality directly, by letting raters build mastery instead of re-entering context repeatedly.
Fast feedback prevents systemic drift
A 24-hour RCA loop keeps a misinterpretation from becoming a pattern.
Put experienced raters on volatile queues
Assigning highly calibrated agents to depleting queues protected quality precisely when conditions were least stable.
Why It Mattered
The programme produced an evaluation pipeline that balanced creative latitude against factual grounding, improved persona alignment, tone consistency and conversational realism, cut annotation latency, and fed concrete insight back into tooling and process decisions. The measurable outcome on the product side was higher conversation authenticity, less verbosity and better persona engagement.
Where This Sits in Hugo’s Evaluation Stack
Persona evaluation is judgment work across factual grounding, tone and safety. It is part of Omni Evals, Hugo’s human evaluation practice for frontier models.
Build your Dream Team
Ask about our 30 day free trial. Grow faster with Hugo!