Research memo · Roux · 9 July 2026
Three parallel investigations — 112 claims extracted, 25 adversarially verified, 0 killed — into whether text-to-image plus image-to-video earns a place in an Awwwards-grade build, and whether Claude can be trusted to judge the result.
Verdict on the central claim
The generative-asset leg survived the parts of the gauntlet I expected to kill it. Awwwards has no AI policy to fall foul of. Seamless looping is solved with a real API primitive. Google will contractually indemnify you against output-infringement claims. Each of those was a live prediction of failure, and each was wrong.
It fails on two things instead, and neither is aesthetic. Raw generated output is uncopyrightable in the US, so you cannot cleanly warrant title in a client contract. And a model cannot be the taste gate — that isn't caution, it's a measured, peer-reviewed result.
Retracted — this report originally asserted a claim it could not support
The first version of this memo stated that the referenced video “demonstrates Claude writing Three.js and shader code” and that “there is no text-to-image step and no image-to-video step anywhere in it.” That claim is withdrawn. It was built from a chapter list and a text-extraction proxy — evidence that cannot establish an absence. The underlying research had already flagged the video as “never independently characterized”; the synthesis asserted it anyway.
Current state: unresolved. Roughly 25 retrieval attempts (timedtext endpoints, innertube, transcript mirrors, Invidious and Piped instances, proxies) failed to obtain a transcript — YouTube’s anti-bot blocks this host. The operator, who watched it, reports seeing a generative model produce a clip. That is the strongest evidence available and it contradicts the retracted claim.
A subsequent agent argued the reverse — “deployed to GitHub/Vercel, therefore the motion must be code.” That inference is invalid and is not relied on here: Vercel serves an MP4 from /public as happily as it serves a shader. It is the same error in the opposite direction.
What is established: the generated-asset pipeline is real and practiced by this creator, in adjacent videos that name the tools outright — “Claude Code + Kling 3.0”, “Claude + Higgsfield” — alongside a documented hybrid of Runway Gen-4 or Kling for background clips plus GSAP and Three.js for scroll motion. Whether this particular video is that pipeline or its code-only sibling remains open.
Separately and independently verified: Claude Design shipped 17 April 2026 as an Anthropic Labs research preview on Opus 4.7, produces code-powered prototypes exportable to HTML, and hands off to Claude Code. Anthropic ships no text-to-3D asset model. That is a fact about the product, not about the video.
anthropic.com/news/claude-design-anthropic-labs · product claim: high, unanimous 3–0 · video content: UNRESOLVED
Nothing below depends on the video. The two hard limits — Thaler and the mean-reversion result — are facts about copyright law and about generate-and-critique loops. They hold whichever pipeline the tutorial demonstrates. If anything, a tutorial shipping a generated video hero would illustrate the WCAG problem documented further down rather than escape it.
Scoring is Design 40%, Usability 30%, Creativity 20%, Content 10%. The official evaluation page contains no AI-generated-asset policy, no disclosure requirement, and no originality-verification rule. Only a reactive plagiarism-report inbox exists. Juries are not instructed to flag, let alone penalize, generated imagery.
The honest caveat: this is an argument from public silence. Private jury instructions are unobservable, and FWA, Webby, and CSSDA policies were never separately confirmed. Absence of a written rule is not evidence that a juror won’t notice.
awwwards.com/about-evaluation/
I told you earlier this was probably the dealbreaker — that no model has a loop primitive and a hero that visibly cuts is unusable. That was my strongest prediction and it is false. Luma’s Ray3 / Ray3.14 expose loop: true on 5-second image-to-video, plus start-and-end keyframe conditioning (frame0/frame1, up to 16 keyframes). Confirmed across Luma’s own docs, Replicate, fal.ai, and AWS Bedrock. Being API-accessible, it is orchestrable and therefore repeatable.
Bounds: looping works only on the 5-second image-to-video path (10-second is text-to-video, no loop, no keyframes) and not in Draft Mode. “Seamless” is a vendor adjective — no independent frame-level benchmark proves zero discontinuity. Spot-check every asset.
lumalabs.ai/learning-center · docs.lumalabs.ai · replicate.com/luma/ray-3.2
Google’s two-layer indemnity covers allegations that generated output infringes third-party IP, not merely training-data claims. Imagen and Veo are on the current indemnified list. But coverage applies only to generally-available model versions accessed through the enterprise Vertex / Agent Platform API, and is conditional on responsible-use conduct.
The consumer Gemini Developer API path — which is what a GEMINI_API_KEY in /etc/imagegen.env almost certainly is — falls under separate terms with no such indemnity. Preview and experimental model versions are excluded too. The enumerated list has already changed once (October 2023 predated Veo entirely), so treat any snapshot as perishable and re-verify at contract signing.
cloud.google.com/terms/generative-ai-indemnified-services · ai.google.dev/gemini-api/terms
Purely AI-generated material cannot be registered for US copyright: the Copyright Act “requires all eligible work to be authored in the first instance by a human being” — Thaler v. Perlmutter, D.C. Circuit, 18 March 2025. Works incorporating AI material are protectable only for the human’s own creative contributions, and applicants must identify and disclaim the AI-generated portions when registering.
For an agency this is concrete, not academic. A standard client MSA asks you to warrant originality, non-infringement, and title. Your human art direction, compositing, and integration are copyrightable. The raw Veo frame is not. Those warranties need explicit carve-outs, and the human authorship needs to be substantive rather than nominal.
media.cadc.uscourts.gov/opinions/docs/2025/03/23-5233.pdf · copyright.gov/ai/ · 88 FR 16190
Hintze, Åström & Schossau (Cell Patterns, Dec 2025) ran autonomous SDXL↔LLaVA generate-and-critique loops: 700 trajectories × 7 temperatures × 100 iterations. All of them collapsed onto roughly 12 dominant “commercially safe” motifs — “generic, commercially viable imagery lacking novelty and surprise.” A second peer-reviewed study found the homogenizing effect persisted after prompt and parameter modification. Temperature tuning does not close the diversity gap.
On the judging side, GPT-4’s self-preference bias measures 0.520 on the Equal-Opportunity metric. LLM judges over-reward low-perplexity, familiar, fluent output regardless of who wrote it — they conflate familiarity with quality. The root cause is contested (perplexity vs. self-recognition, a 2–1 split among verifiers), but the bias itself is not.
cell.com/patterns/fulltext/S2666-3899(25)00299-5 · arXiv:2410.21819 · arXiv:2404.13076
Current generative architectures may be fundamentally limited in genuine creative exploration when operating autonomously… human–AI collaboration may be essential to preserve variety and surprise. Hintze, Åström & Schossau — Cell Patterns, 2025
Read that against what practitioners independently report about Claude Design: outputs converging on a house style — teal gradients, a serif headline, a blinking status dot, cards nested inside cards — because the underlying design skill falls back to a small preset set when the prompt is loose. Two investigations, two pipelines, the same defect. Mean-reversion under weak steering. The mitigation is identical in both and has nothing to do with model choice: inject strong, specific, external art direction, or the tool hands you its prior.
| Axis | Claude → Three.js / GSAP | T2I → I2V → video |
|---|---|---|
| Reduced motion | Free — guard the timeline, skip the RAF loop | Manual — browsers do not honor prefers-reduced-motion for <video autoplay>; open WHATWG issue. Ship a second asset and swap. |
| WCAG conformance | Straightforward | Needs a visible control — SC 2.2.2 is Level A and demands pause/stop for motion over 5s. An OS preference does not satisfy a page-control requirement. |
| Copyright / title | Clean | Uncopyrightable raw — needs MSA carve-outs |
| Indemnity | N/A | Enterprise path only — Vertex GA, not the consumer API |
| Versionable / diffable | Yes — it’s source | No — opaque pixels, reroll and pray |
| Marginal cost | $0 on plan auth | Metered — per image, per second, per client |
| Art-direction ceiling | Bounded by what you can code | Photoreal scenes you couldn’t afford to build |
| Used by top studios | Documented — Immersive Garden, Active Theory, Locomotive, darkroom | No studio case study found — though the tutorial/creator space demonstrably ships it (Kling, Runway Gen-4, Higgsfield). Cannot distinguish “rare at the top” from “undisclosed.” |
Every named-studio case study describes real-time WebGL and shader code, not generated video. Immersive Garden’s Havas SOTM: Blender models exported to Three.js, Perlin-noise shader transitions. Active Theory runs a bespoke WebGL engine and renders text in WebGL to dodge DOM compositing jank. Not one studio is on record shipping a text-to-video hero — and the agent was appropriately careful to say it cannot tell whether that means rare or merely undisclosed.
Coca-Cola’s AI Christmas ads drew sustained backlash two years running; positive sentiment fell from 23.8% to 10.2%. McDonald’s Netherlands and Valentino both pulled AI campaigns. A 2025 Nuremberg Institute study found disclosed AI authorship carries a measurable trust penalty — lower perceived naturalness, weaker purchase intent.
The search for a credible art director publicly defending AI-generated hero visuals for premium brand work returned nobody on record. I want to flag the obvious confound rather than lean on the result: the strongest pro-craft source found was Getty, which sells stock photography and has a direct commercial interest in the conclusion. A one-sided literature is weaker evidence than it looks.
The route that inherits code-motion’s safety while buying generated art direction: use the generated frame as a texture, not as the motion. A generated still feeds a shader-driven Three.js scene. Animation stays in code, so reduced-motion and perf gating remain free, while the art direction comes from a model. Be clear-eyed that this is architecturally sound but empirically unproven — no named award-winning site is documented shipping exactly this.
cinematic-3d-render already produces deterministic, re-renderable, fully-owned frames. Generation earns its place only where Blender cannot reach — photoreal scenes you’d otherwise need a 3D artist and a week to build.I am not going to invent numbers to fill this table. The research resolved licensing structure but did not resolve per-second pricing for Veo 3, Kling, Runway, or Sora — that leg of the brief was dropped under budget. Treat everything below as verify-at-use.
| Item | Figure | Provenance |
|---|---|---|
| Imagen 4 | ~$0.04 / image | Your own prior note, not re-verified here |
| Veo 2 | ~$0.35 / second | Your own prior note, not re-verified here |
| FLUX “Builder” | 10K img/mo, 1 domain | bfl.ai/licensing — explicitly “not meant for client use” |
| FLUX “Professional” | 100K img/mo, 3 domains | First 3 named clients included; per-client fees beyond |
| Veo 3 / Kling / Runway / Sora | Unresolved | Not established by this research |
| Claude, all legs | $0 marginal | Max plan auth |
The FLUX finding matters more than it looks: the free and entry tiers do not license agency client deliverables at all. Client work is a paid, per-client-metered arrangement. That is a licensing gate, not a price — and it is the kind of thing that surfaces after you’ve already shipped.
prefers-reduced-motion only as something to add, never as something tested.