A persona chat conversation on a phone beside an evaluation panel comparing two candidate responses on factuality, preference, tone and social memory

How Hugo Strengthened Persona-Driven Conversational AI at Scale

Author: Hugo

A persona-led chatbot has a harder job than a general assistant. It has to be accurate, and it has to stay in character, and when those two pull in opposite directions something has to give. Deciding which should give, across millions of conversations, is a human judgment problem.

A technology organization running a creator platform for persona-led conversational agents engaged Hugo to evaluate and strengthen the conversational quality of those personas. Across eight workflows the team delivered ~1.16 million annotations, reached 97.76% QA against a 90% target and 92.5% accuracy against 81.1%, at an average handling time 27% faster than target.

Why It Was Hard

Two workflows that look identical and are not

The sharpest challenge was structural. Preference and Factuality share the same inputs, the same layout and nearly the same steps: user prompt, bio and style directives, social memory facts, message history, model responses. What differs is the lens. Preference asks which response is better. Factuality asks which is true. A rater running the wrong lens produces confident, well-formed, useless data.

The Preference workflow, step one, showing responses assessed against social memory facts
Preference, step one: responses checked against social memory.
The Factuality workflow, step one, showing prompt evaluation with the same inputs as the Preference workflow
Factuality, step one. Same inputs, different question.
The Preference workflow style evaluation step
Preference then evaluates style against the persona’s directives.
The Factuality workflow social memory evaluation step
Factuality instead evaluates the prompt against social memory.
The Preference workflow conversationality evaluation step
Preference: conversationality.
The Preference workflow final response ranking step
Preference closes on a ranking.
The Factuality workflow non-factual criticality evaluation step
Factuality closes on criticality: how much the inaccuracy actually matters.

Multi-step reasoning over full histories

Ranking a response meant reading the whole conversation, categorising user intent from it, and only then judging the candidates. Context could not be skipped.

A full conversation history under review, used to categorise user intent before ranking responses
The unit of work was the conversation, not the message.

Efficiency that moved for reasons outside the work

Seed data shortages, tooling timeouts, skipped complex jobs and network instability all pulled at occupancy and non-handling time.

Rater harmony

With this density of subjective decisions, agreement between raters, QAs and client auditors is not a reporting metric. It is the product.

Hugo
Alignment tracked as a first-class output.

How Hugo Approached It

Upfront preparation

Workflow-specific SOPs, scoring rubrics, calibration packs and simulation exercises, built around persona depth, tone, preference logic and factual accuracy. A 24-hour feedback and RCA loop closed client-identified issues before they could recur.

Team structure, organised against cognitive load

176 raters and 43 QPEs were organised into domain-aligned units so context mastery was preserved rather than diluted. Frequent queue switching was replaced with a main plus fallback queue structure, and every rater was cross-skilled on exactly one fallback: flexibility without losing specialisation.

Four dedicated QA coaches ran targeted calibration pods at a 1:15 ratio, clarifying ambiguous cases and deepening mastery in the workflows where misinterpretation was most likely. Small pods outperformed broad training calls consistently.

A triple-layered review loop

  • Rater review. Each morning, raters worked through the prior day’s audits, identified patterns and shared what they found.
  • QA review. QAs validated consistency and surfaced the nuanced cases that guidelines had not yet resolved.
  • Client QA reconciliation. Flagged discrepancies were reviewed inside 24 hours, so recalibration happened while the context was still fresh.

Daily audit-of-audit reviews reinforced scoring discipline and held QA above 97% from July 2024 onward.

Continuous calibration

Daily RCA ran across quality, utilization, occupancy and handling time to catch trends early. Audit-of-audit sessions were used specifically to refine the subjective rubrics and settle the Preference-versus-Factuality boundary. A structured communication cadence with the client kept tool and instruction changes from sitting unactioned.

Efficiency and reliability

New queues were tested early to catch timer and tooling errors before they hit production, best practice for error-prone jobs was reinforced, and escalation pathways were shortened to cut downtime.

Results

  • ~1.16 million annotations across eight workflows.
  • 97.76% QA, exceeding target by 7.76 points.
  • 92.5% accuracy, exceeding goal by 11.4 points.
  • Average handling time 1,428 seconds against a 1,974-second target, 27% faster.
  • 97.75% occupancy with non-handling time at 1.96% after optimisation.
  • A 43-week sustained delivery cadence.

What We Learned

Similar-looking workflows need explicit separation

Where two workflows share structure but differ in lens, calibration and worked examples are the only reliable fix. Rubric text alone will not hold the line.

Small coaching pods beat broad training

Nuance transfers when feedback is contextual and tailored. It does not transfer well on a group call.

Stability lowers cognitive load

Stabilising queue assignments improved quality directly, by letting raters build mastery instead of re-entering context repeatedly.

Fast feedback prevents systemic drift

A 24-hour RCA loop keeps a misinterpretation from becoming a pattern.

Put experienced raters on volatile queues

Assigning highly calibrated agents to depleting queues protected quality precisely when conditions were least stable.

Why It Mattered

The programme produced an evaluation pipeline that balanced creative latitude against factual grounding, improved persona alignment, tone consistency and conversational realism, cut annotation latency, and fed concrete insight back into tooling and process decisions. The measurable outcome on the product side was higher conversation authenticity, less verbosity and better persona engagement.

Where This Sits in Hugo’s Evaluation Stack

Persona evaluation is judgment work across factual grounding, tone and safety. It is part of Omni Evals, Hugo’s human evaluation practice for frontier models.

Build your Dream Team

Ask about our 30 day free trial. Grow faster with Hugo!

Share