How Hugo Delivered Reliable Human Evaluation for AI-generated Game-world Videos
Generating a synthetic game-world video is now the easy part. Judging whether the result is any good is not.
A technology organization building agentic AI systems had a pipeline that could assemble game-world video at volume: real game assets and gameplay captures stitched into coherent scenes, with AI-generated music laid over the top. What it could not do was tell, reliably, whether any given output was fit to use. Hugo was engaged to supply that judgment. Over the run of the Horizon AI Weels project the team delivered 185,864 evaluations, at an average handling time of 3 minutes 59 seconds against a 5-minute target, and moved client-audited quality from 75% to 91.21%.
Five Dimensions, Judged Simultaneously
Every video was assessed across five axes at once:
- PII. Real names, visible usernames, nametags floating above avatars.
- Visual quality. Hallucinated elements, wrong orientation, fake interactive overlays, poor transitions, shaky footage.
- Audio quality. Hallucinated or distorted sound, robotic voices, cut-off dialogue, audio drifting out of sync with picture.
- Storyline. Whether scenes cohere and the narrative actually progresses.
- Accuracy. Correctness of text, numbers and gameplay mechanics, and fidelity of environments, character models and avatar assets.
The distinguishing demand was domain knowledge. Raters were not reviewing video in the abstract. They were deciding whether an AI system had faithfully reproduced a specific real game environment, which requires knowing what that environment actually looks like and how it behaves.
Why It Was Hard
The interface changed underneath the work
Partway through, the evaluation tool lost its hard-stop response gates. Those gates had capped how many error types a rater needed to consider once a particular flag was triggered. With them gone, the full scope of dimensions opened on every job regardless of earlier selections. Overnight, a scoped task became a much broader one, and it exposed misalignments that the narrower flow had been hiding. Accuracy dipped until the team recalibrated.
Neighbouring error categories
The framework asked raters to separate error types that look alike but mean very different things to a model team. Audio hallucinations against merely robotic voices. Wrong content orientation against incorrect aspect ratio. Fake interactive buttons, hallucinated by the model, against real interface elements.
The most persistent boundary was visual hallucinations versus gameplay inaccuracies. A hallucination is content with no basis in the real game at all: an object, environment or character that does not exist. A gameplay inaccuracy shows real game elements in the wrong context, sequence or configuration. Conflating the two produces an evaluation signal a research team cannot act on, because the fix for each is different.
How Hugo Approached It
Game-world familiarization, before any scoring
Raters played the games. Guided exploration of the real environments was built into onboarding and sustained through the project, with client-provided access protocols for each world and a standing requirement to explore any unfamiliar game before evaluating trailers from it. Judgments were anchored in experience of the environment rather than surface inference from the clip.
A living escalation log
Edge cases were frequent enough to need their own routine. Each morning before queues opened, raters and QPEs reviewed newly escalated cases and prior resolutions, worked through commonly misjudged boundaries in short coaching sessions, and logged ambiguous jobs for client review. Client rulings were folded back into the calibration materials, so the reference guide grew more complete as the project ran.
Dual-cadence quality assurance
Daily internal audits caught individual drift; weekly client reviews caught systemic interpretation gaps. Raters below threshold got dimension-level feedback before re-entering live queues, and client feedback landed in the next day’s calibration cycle rather than the next week’s.
Certification as a gate, not a formality
Production access required 80% on a Project Knowledge Test, following structured training across all five dimensions. Purpose-built calibration materials targeted the known-hard distinctions, particularly hallucination against inaccuracy, using annotated real cases to mark where one category ends and the next begins.
The Evaluation Workflow
Each video moved through the same sequence.






Results
- 185,864+ annotations across all production queues.
- Client-audited quality from 75% to 91.21%, a 16.21 point improvement, achieved while guidelines shifted, the UI changed and evaluation scope expanded.
- Average handling time of 3 minutes 59 seconds against a 5-minute target.
- Rater pool scaled from 94 in peak production down to 19 billable raters as the work stabilised.
The shape of that quality curve matters as much as its endpoint. It rose steadily as familiarity deepened, rather than peaking early and drifting.
What We Learned
Domain familiarity is a prerequisite, not a nice-to-have
Raters who had navigated the actual game environments produced measurably more consistent accuracy scores than those working from surface judgment. For content that reproduces a real world, you cannot evaluate the reproduction without knowing the original.
Adjacent error categories need their own materials
Where two categories share surface features, a guideline document is not enough. Annotated examples that mark the boundary are what reduce repeat misclassification.
A tooling change is a calibration event
Removing the hard stops was a UI decision with a direct quality cost. Treating it as a formal recalibration trigger, with dedicated sessions and heightened drift monitoring, let the team absorb it without sustained degradation. Tooling changes on an evaluation programme are never just tooling changes.
Two cadences beat one
Daily internal audits caught individual drift fast. Weekly client reviews caught systemic gaps. Neither alone would have held the line across an expanding scope.
Why It Mattered
The client got a dependable quality signal across five dimensions on one of its most complex generative systems, and a rater pool whose judgment sharpened over the life of the project. Evaluating synthetic video against a real, specific world is a domain-knowledge problem wearing the clothes of an annotation problem. Staffing it accordingly is the whole game.
Where This Sits in Hugo’s Evaluation Stack
Horizon AI Weels was multimodal evaluation of generated video across visual, audio and narrative dimensions. It is part of Omni Evals, Hugo’s human evaluation practice for frontier models.
Build your Dream Team
Ask about our 30 day free trial. Grow faster with Hugo!