Six cases
What the loop looks like in practice
Each card below is one real generation from the run above. Watch the video on the left through
the customer’s eyes — what they actually saw and clicked thumbs-up or
-down on. Read the right column through the judge’s eyes — what it scored,
what it explained, and how that explanation changed (or didn’t) when the customer’s feedback was handed to it.
A cyan ring around a dimension chip means the score moved between the first pass and the feedback turn. That is
the per-case progression matrix: same judge, different evidence, different verdict.
▲ Customer upvote
gmail.com
1 → 1 (held)
When the judge was right — about the wrong brand
The workspace says gmail.com, but only because the customer signed up with a personal Gmail address — brand inference defaulted to their email domain. They actually wanted a Roblox video, uploaded their own Roblox assets, and were delighted with the result. The generator tried to align it to the Gmail “brand,” and the judge correctly dinged it for “focusing on Roblox rather than Gmail.”
What the customer asked for
here is my idea for a script: 🎬 VIDEO SCRIPT (2 MINUTES)
⏱️ 0:00 – 0:15 (Hook)
VO:
“Roblox just hit one of its biggest turning points ever… after years of explosive growth, the player count is starting to drop from its…
What the customer said
great pacinggood scripton brandsmooth animationsgood voiceover
“can you add subscribe and like for more at the end of the video and have a roblox player just walking around in a random game in the backkround.”
overall 1first pass · judge sees prompt + brand evidence + frames
“The video fails the brief by focusing on Roblox rather than Gmail, ignoring key brand assets and demos, and running under the requested 2-minute duration, making it a dealbreaker.”
per-dimension
2ast1prd1dmo1brd1str3exc1ovr
overall 1after the feedback turn · judge re-reads with the customer’s note
“Feedback assessment: unchanged - The user’s positive reaction and request to add a subscribe/like call-to-action and a Roblox character do not address the fundamental misalignment with the brief (wrong product focus, missing Gmail assets, incorrect duration), so the original rating stands.”
per-dimension (cyan ring = changed)
2ast1prd1dmo1brd1str3exc1ovr
Why it matters · The judge’s score is faithful to the rubric — but the rubric was built against a brand the customer never claimed. This is the auditable trace doing its real job: surfacing an upstream pipeline bug (brand inference is brittle on personal-email signups) that no headline metric would ever catch. The customer’s upvote was correct; our system was wrong about who they were.
▲ Customer upvote
medilaunch.health
2 → 2 (held)
When the judge held the line — substrate intact
Real brand, real assets, four positive reasons ticked, and a constructive note (“needs more animation — bit faster pace”). The judge held at 2 anyway, because the video missed multiple hard constraints: length, required asset usage, and prescribed scene content. The feedback turn didn’t budge it.
What the customer asked for
Create a 45-second product explainer for MediLaunch, a private-label telehealth
platform. Tone: confident, clinical, premium. Pace: brisk but not rushed.
VISUAL DIRECTION
- Brand color: #2F6FE8 (calm clinical blue). Ac…
What the customer said
on brandsmooth animationsgood voiceovergreat pacing
“needs more animation - bit faster pace... more info”
overall 2first pass · judge sees prompt + brand evidence + frames
“A polished draft visually, but it misses hard constraints—length, required asset usage, and key scene content—resulting in failure to meet the core brief.”
per-dimension
2ast4prd2dmo2brd2str3exc2ovr
overall 2after the feedback turn · judge re-reads with the customer’s note
“Feedback assessment: unchanged - the user’s comments cheer the animations and pacing but do not address the critical brief failures around length, asset usage, and scene content that drive the low score.”
per-dimension (cyan ring = changed)
2ast4prd2dmo2brd2str3exc2ovr
Why it matters · Positive customer feedback is real signal — it just isn’t the whole signal. The judge’s job is to anchor on what was in the brief that wasn’t done, and let the loop expose the disagreement so a human can decide whether the rubric is too strict or the customer is too generous.
▼ Customer downvote
pwc.com
4 → 3
When the customer’s ears caught what the judge’s eyes can’t
The judge sees rendered frames — stills. Audio drift is invisible to it. The customer flagged out-of-sync narration; the judge accepted the evidence and lowered the score.
What the customer asked for
You are a PwC Investigations and Forensics Partner with over 50 years of experience. You’ve developed an AI Enabled Investigations Workbench which supports the end-to-case investigative lifecycle, from case intake throug…
What the customer said
“The voice changed during the video. The video narration was out of sink with the image / video.”
overall 4first pass · judge sees prompt + brand evidence + frames
“A strong, demo-first presentation with excellent fidelity to assets and product details; minor tweaks to typography and interstitial styling would bring it fully on-brand.”
per-dimension
5ast5prd5dmo4brd5str4exc4ovr
overall 3after the feedback turn · judge re-reads with the customer’s note
“Feedback assessment: changed - the reported audio/visual synchronization and voice inconsistency issues reveal a significant executional usability problem, lowering the overall score to a 3.”
per-dimension (cyan ring = changed)
5ast5prd5dmo4brd5str3exc3ovr
Why it matters · Frames aren’t the whole video. The feedback turn is how the loop reaches into the parts of the artifact the judge can’t observe on its own.
▼ Customer downvote
helloinvoice.co
4 → 3
Pacing is another thing stills can’t show
“The video was meant to be upbeat and energetic, but this feels slow and lethargic.” The judge’s first pass called it polished; after reading the note, it agreed the energy was off.
What the customer asked for
Create a product launch video announcing two new features: Work, Projects.
Start with an opening logo.
Then open with a text hook that captures the problem users face today. Then introduce using some broll that is relate…
What the customer said
too slowvoice quality
“The video was meant to be upbeat and energetic, but this video feels slow and lethargic”
overall 4first pass · judge sees prompt + brand evidence + frames
“This is a solid, demo-focused video that hits the feature brief and uses real HelloInvoice assets, but generic gradient styling and watermarks keep it from feeling fully polished in brand voice.”
per-dimension
5ast5prd5dmo4brd5str4exc4ovr
overall 3after the feedback turn · judge re-reads with the customer’s note
“Feedback assessment: changed - User feedback on the slow, lethargic pacing and voice quality reveals underweighted issues in tone and energy, warranting a lower overall score for notable execution problems.”
per-dimension (cyan ring = changed)
5ast5prd5dmo3brd5str3exc3ovr
Why it matters · Two cases like this in a row, same root cause: the judge’s evidence is incomplete by construction. The loop is how we patch that without paying for full video understanding on every gen.
▼ Customer downvote
www.emptrackai.online
overall 1
When the judge already knew exactly what was wrong
Customer ticked “feels generic.” The judge had given it a 1 on first pass — and walked the rubric clean through every dimension to show why.
What the customer asked for
Create a 10-second text animation overlay. Export with transparent background.
The animation shows 3 problem keywords, one at a time, each
slamming in and then fading out. light cinematic style.
Timing and words:
0:0…
What the customer said
feels generic
overall 1first pass · judge sees prompt + brand evidence + frames
“This submission misses multiple core brief requirements—transparent background, brand assets, product references, and demo content—representing a fundamental mismatch with the user’s prompt and brand guidelines.”
per-dimension
1ast1prd1dmo1brd1str3exc1ovr
Why it matters · When the loop confirms a downvote with concrete rubric language, that row is a calibration win — both eval system and human agree, and we have a regression test for the future.
▼ Customer downvote
numeo.ai
5 → 5 (held)
When the judge is probably wrong — and the loop says so
The customer downvoted with “visual mistakes.” The judge gave a 5 and held there, calling the note “too vague to override the rubric.” That’s exactly the kind of disagreement a human should look at.
What the customer asked for
Numeo Load Hub is a unified loadboard for freight dispatchers. Instead of jumping between DAT, Truckstop, RXO, TQL, C.H. Robinson, J.B. Hunt, Schneider, Arrive, Echo, Emerge, Parade, and 5+ other boards, dispatchers sear…
What the customer said
other (negative)
“video is a bit off, with many visual mistakes in it”
overall 5first pass · judge sees prompt + brand evidence + frames
“This video fully satisfies the brief: it leverages the provided assets and script, demonstrates Load Hub features concretely, and embodies Numeo’s brand tone and visual style end-to-end.”
per-dimension
5ast5prd5dmo5brd5str5exc5ovr
overall 5after the feedback turn · judge re-reads with the customer’s note
“Feedback assessment: unchanged - The user’s comment about “visual mistakes” is too vague and subjective to indicate any concrete failure against the rubric’s criteria, and the video remains fully compliant and on-brief.”
per-dimension (cyan ring = changed)
5ast5prd5dmo5brd5str5exc5ovr
Why it matters · The point of the loop isn’t that the judge is right — it’s that every score is auditable. A 5-vs-downvote standoff with a clean trace is a row queued for human review, not noise to argue about.
▼ Customer downvote
smarterpay.com
overall 2
When the judge catches the spec violation directly
Customer note was just “No images.” The judge identified the deeper miss on its own — the brief asked for 30 seconds, the video ran long, and the scene content didn’t match what was prescribed.
What the customer asked for
[Scene 1 – Problem | 0–5 sec]
Visual: Someone manually uploading files, checking folders, looking frustrated.
Voiceover:
“Still manually uploading payment files? Managing folders, encryption, and submissions can slow you…
What the customer said
“No images”
overall 2first pass · judge sees prompt + brand evidence + frames
“By significantly missing the 30-second duration, ignoring prescribed scene content, and failing to show the actual product, this video does not meet the core requirements of the brief.”
per-dimension
2ast4prd2dmo3brd2str3exc2ovr
Why it matters · Structural adherence (length, scene order, required assets) is the dimension the judge can reliably enforce from the evidence packet alone — this is where the judge earns its keep without needing the customer’s help.