A performer in a motion capture suit with a skeletal mesh overlay beside an evaluation panel for motion caption, asset QA and audio QA

How Hugo Scaled High-Reliability Annotation Across Eight XR Workflows

Author: Hugo

Avatars in virtual and augmented reality only feel right when a lot of separate things are right at once: motion, facial expression, audio clarity, eye tracking, segmentation, and how a stylized human is classified. Each of those is a different evaluation problem, with a different failure mode and a different kind of expert judgment behind it.

A technology organization building the AI systems behind realistic XR avatars engaged Hugo to annotate and evaluate across the full set. The programme delivered ~3.49 million annotations across eight workflows at 93.4% accuracy against an 85% goal, with average handling time 34.7% faster than target.

Eight Workflows, Four Modalities

The work spanned image, video, audio and stylized content, each with its own rules.

Hugo
Motion captioning. Free-form description of human action and behaviour in video, feeding a captioning model.
Hugo

Video rule-based tagging. Posture, movement, mood and object interaction for a highlighted person across a 20-second segment.

Hugo

Asset QA. Skeletal mesh labelled success or error across 5-second videos.

Hugo
Audio QA. Two models compared on naturalness and fidelity to the source audio.
Hugo

Skin tone evaluation. Benchmarking model performance across skin tones using grouped categories, so coverage gaps are measurable rather than assumed.

Hugo

Image classification. Anatomical correctness of stylized characters, flagging issues and rejecting multi-person images.

Hugo
OCT whole eye QA. Motion artifacts and retina visibility in OCT eye imaging, the most technically specialised queue on the programme.

Why It Was Hard

Free-form description creates legitimate disagreement

Motion captioning asks for prose, and prose varies. One annotator writes “the man raises his right hand and waves.” Another writes “the man lifts his right hand and moves it side to side in a greeting gesture.” Both are correct. Under QA review, that difference can read as inconsistency when it is really vocabulary.

Hugo
Same action, two valid descriptions. Resolving this is a rubric problem, not a rater problem.

Visual and motion complexity

Segmentation work required assigning up to 40 labels to stylized images with blurred edges, blended colours and deliberately distorted anatomy. Motion captioning meant capturing facial expression, body orientation, camera movement and environmental context together.

A stylized image with blurred edges and blended colours requiring up to forty segmentation labels
Stylized source material, where the edges themselves are ambiguous.

Modality switching against moving guidelines

Annotators moved between image, video, audio and stylized queues, each with distinct rules, while guidelines were updated frequently enough to require ongoing retraining and rapid realignment.

One of several distinct annotation interfaces raters moved between across modalities
Each modality carried its own interface and its own conventions.

Consistency across 180+ annotators

Uniform interpretation across segmentation, classification and motion captioning at that headcount is only achievable with deliberate calibration and tight feedback loops.

How Hugo Approached It

Upfront preparation

Workflow-specific SOPs, annotation guides, learning materials and assessments were built for each queue. Internal tests acted as learning checkpoints. Quick reference guides were written for the hardest workflows, Motion Prior and Motion Captioning, so raters could resolve difficult cases without stalling. An enhanced quality process ran a daily RCA loop that addressed all client feedback within 24 hours.

Team structure matched to the work

  • Generalist raters handled primary annotation across motion captioning, motion prior, skin tone labelling and stylized anatomy.
  • STEM-skilled raters took the technically demanding queues, notably OCT whole eye evaluation and advanced anatomy review, identifying motion artifacts and subtle structural inconsistencies.
  • QA ran internal audits, enforced rubric adherence and tracked error patterns.
  • Training owned onboarding, certification, guideline walkthroughs and recalibration driven by QA findings.
  • Management held workforce planning, escalations, delivery timelines and client coordination.

A multi-layered evaluation framework

  • Rater review, daily. Annotators reviewed their prior-day audited jobs, identified interpretation gaps and flagged edge cases and disagreements early.
  • QA review. Validation of accuracy plus active surfacing of ambiguity patterns for clarification.
  • Client QA reconciliation. Flagged discrepancies resolved on a 24-hour RCA loop, enabling recalibration before drift could set.

Results

  • ~3.49 million annotations across eight workflows.
  • 93.4% accuracy, surpassing the goal by 8.4 points.
  • 85.3% utilization, 0.3 points above target.
  • Average handling time 1,289 seconds against a 1,974-second target, 34.7% faster.

What We Learned

Precision compounds

XR workflows turn on micro-detail: pixel edges, motion sequencing, audio clarity. All three improved with stronger visual examples and calibration pods.

Description quality improves with a writing framework

On Motion Prior, grammar and sentence structure directly affected annotation quality. Where the deliverable is prose, teaching the prose is part of teaching the task.

Modality switching needs pacing

Annotators stabilised fastest when allowed to master one workflow before moving to the next. Parallel exposure across four modalities slowed everyone down.

RCA accelerates recovery

Daily RCA with error-pattern tracking and explicit next-day corrections reduced drift across both segmentation and motion tasks.

Tool reliability is a quality variable

Job loading failures and thin queue volume measurably hurt efficiency. Early escalation improved stability.

Pattern tracking closes the loop

Documenting reasoning pitfalls, prompt ambiguities and interpretation gaps created a direct line from what raters saw to how guidelines were written. That feedback path is what turns an annotation vendor into a research input.

Why It Mattered

The programme produced more accurate and consistent visual and motion datasets, improved segmentation precision on stylized image sets, sharpened motion understanding for expressive avatars, raised classification accuracy across stylization and quality tiers, and fed concrete guideline and tooling improvements back to the client’s research teams.

Where This Sits in Hugo’s Evaluation Stack

XRCIA spanned perception, motion, audio and specialised medical imaging queues. It is part of Omni Evals, Hugo’s human evaluation practice for frontier models.

Build your Dream Team

Ask about our 30 day free trial. Grow faster with Hugo!

Share