WB Games GenAI Creative Specialist — Take-Home Assessment
Deliverable 5 from the brief, “A process breakdown.” Documents tools, prompts, iterations, guardrail handling, and time spent.
1. The brief
Brief: Show a dragon’s progression from egg (Day 1) to war beast (Day 100) in the Game of Thrones: Dragonfire world. A 30-59s vertical UA video and a 9:16 static, generated only with Kling 3.0 (video) and Nano Banana 2 / Nano Banana Pro / GPT Image 2 (stills). Hard guardrail: the dragon can never show separate arms at any growth stage.
The concept, in one sentence: A dragonrider’s gauntleted hand is the one thing that never changes while the dragon under it goes from a nipping hatchling to a war beast, Day 1 to Day 100.
Progression is read through the hand’s relationship to the dragon, not a montage of bigger dragons: care on Day 1 becomes command on Day 100.
2. The final deliverables
Final video
Final static

Before: the raw GenAI static, pre-comp

The before/after is worth keeping visible. It’s the fastest way to show what the AI generated versus what got built on top of it in Photoshop, which is exactly the “AI as a tool, not the output” distinction this brief is testing for.
3. Platform & sound design hypothesis
Platform: Meta (Reels / Stories / Feed), not TikTok. TikTok’s audience skews toward UGC/creator-led video, and GenAI video underperforms there; Meta rewards polished, cinematic motion and rewards silent readability, which this concept is built around.
Sound-off / sound-on design: the DAY counter is the sound-off spine: same position, same type, cuts in with the shot, never animates. Sound-on is rewarded, not required: a hatchling chirp-screech motif, pitch-shifted down through the piece (shot 2 to 3 to 5 to 7), becomes the Day 100 roar at the same contour. That’s an audible progression that still reads at zero volume. Two deliberate silences (the cold open before the crack, and the cut into shot 7) do more work than a wall of sound would.
Music: a licensed stock orchestral/cinematic track. It opens simple and playful under the hatchling beats, builds through the growth beats, and lands its epic swell as the dragon reaches war-beast scale.
4. Workflow
Text to image to video to edit. Never text-to-video on the dragon itself: every video generation is image-to-video from a locked, anatomy-checked still.
- Style master (Nano Banana Pro, multi-panel): locks palette, lighting, camera language, and the gauntlet design once, up front.
- Stage sheets (S1-S5, Nano Banana Pro): five turnarounds, one per growth stage, same anatomy phrase every time. This is the “no arms” evidence, produced before a single video credit is spent.
- Shot stills (Nano Banana Pro): one per shot, built from the style master plus the relevant stage sheet as references.
- Video (Kling 3.0, weavy.ai): 3.0 Pro image-to-video from the locked still for single-subject shots; O3 (or 3.0 Pro with small motion) where the hand and dragon both have to hold together across motion.
- Edit: Premiere for the cut, retime, counter, end card; After Effects for artifact masking and the end card composite; Topaz for upscaling after picture lock; Audition for the motif build and final mix.
The full weavy.ai node board is linked further down, alongside the After Effects and Photoshop screenshots.
5. Style master and stage sheets: the “no arms” evidence

Character Sheets





Five images, zero arms, generated before a single video credit was spent. That’s the anatomy guardrail’s actual evidence, produced before a single second of video was rendered.
6. The “wyvern” vocabulary note
Models carry a strong four-legged “dragon” prior and a strong two-legged “wyvern” prior. Naming the anatomy by its own word is more reliable than negating “arms” after the fact. Every dragon prompt in this project carries the same anatomy phrase, verbatim:
wyvern anatomy: two muscular hind legs, leathery bat-like wings that double as forelimbs with a single wing-claw thumb at the wrist, no separate front legs, no arms
And the same Kling negative prompt, every shot:
four legs, front legs, forelegs, arms, hands on the dragon, quadruped, T-rex arms, extra limbs, extra wings, human face, actor likeness, text, watermark, logo
The answer to “how did you handle no-arms” is vocabulary, locked stills, and reference conditioning. It repeats reliably; it isn’t luck.
7. Shot table
| Shot | Still |
|---|---|
| 1 (egg) | S02 |
| 2 (hatch) | S03 |
| 3 (feed) | S04 |
| 4 (first fire) | S05 (GPT Image 2, rearing pose) |
| 5 (liftoff) | S06 |
| 6 (aerial, continuous from 5) | (no separate still: I2V continuation off S06’s final frame) |
| 7 (Day 100 return) | S08 |
| 8 (fire) | S09 |
| End card | S10 |
S05 needed a regeneration pass: the first gen read as quadrupedal, so it was rejected and rebuilt with a rearing pose that put full weight on the hind legs instead.
Kling Standard vs. Pro: started on Kling Standard to keep credit spend down, but its failure rate on this creature ran higher than expected. More rerolls than the savings were worth. Switched to Kling Pro, which needed fewer rerolls to land a usable result despite the higher per-generation cost. Net cheaper in practice, even at a higher sticker price per clip.
8. Sound
- The motif: one hatchling chirp-screech, pitch-shifted down through the piece, becomes the Day 100 roar at the same contour.
- The two silences: the cold open (before the crack) and the cut into shot 7 (Day 80 return) both drop everything.
- Phone-speaker check: The audio is balanced and well-structured for mobile hearing and desktop hearing.
9. Static rationale
The static is a dual-dragon-flight composition: the war-beast-stage dragon (Day 100) banks above the hatchling-stage dragon (Day 1), both over a burning city, matched to the structure of an ad already running in this exact placement. The trade: video and static now share the Day 1 → Day 100 concept and type, but not an identical shot, and not the video’s hand device. That’s a deliberate choice, not an oversight. Worth naming directly in the process breakdown rather than letting a reviewer wonder why the pair doesn’t visually match.
The raw generation vs. the final composited static are both shown above in Section 2. The gap between them is the Photoshop pass: logo lockup, DAY labels, and CTA built and placed by hand, not generated.
10. Tools list
| Tool | What it did |
| Nano Banana Pro | Style master, stage sheets, shot stills |
| GPT Image 2 | S05 regeneration (anatomy fix) |
| Kling 3.0 (Standard, then Pro) | Image-to-video for all 8 shots |
| weavy.ai | Node-based orchestration, prompt variables ([Variable 1]) |
| Adobe Premiere | Cut, retime, counter, final assembly |
| Adobe After Effects | Artifact masking, end card composite |
| Adobe Photoshop | Static image build |
| Adobe Audition | Motif build, final mix |
11. Process assets
weavy.ai board: https://app.weavy.ai/flow/3lixel7O4TQrGuezmdoqYM


12. Time log
Actuals, not the original plan:
| Phase | Time |
| Concepting + stills (Wednesday evening) | ~4–5h |
| Nano Banana / GPT Image 2 passes | ~4h |
| Kling video passes | ~4–6h |
| Process breakdown, packaging | ~4h |
| Total | ~16–19h |
Generation spend: roughly $60 in weavy.ai credits (~72,000 credits), driven mainly by Kling rerolls. See the Standard-vs-Pro note in Section 7.
Worth stating plainly rather than rounding down: Jessica’s framing was “about a day,” and this ran closer to two. That gap is itself an honest data point for the “time spent” ask. Most of it sat in the video generation pass, not the concepting or the edit.
13. What I’d test next
Full POV structure. Several shots already carry a POV read (feed, first fire, aerial continuation). Building the whole video as continuous rider POV, rather than anchoring on the hand as a recurring device, would test whether embodiment reads as more visceral on Meta’s sound-off feed than the current hand-focused approach. Trade-off: loses the hand as a single, nameable device, but opens the door to the voiceover variant below.
Continuous morph structure. Chain the growth stages as one uninterrupted motion instead of 8 discrete shots, with each clip starting and ending on the same framing as the next, so the dragon visibly morphs from hatchling to war beast across a single sustained move rather than cutting scene to scene. Hypothesis: continuous transformation reads as more premium and more differentiated from typical cut-heavy UA pacing, at the cost of much tighter I2V frame-matching discipline across every transition.
Platform-specific voiceover. The POV structure pairs naturally with a synthesized voice track on TikTok, where AI-voice narration reads as native to the platform rather than as a shortcut. On YouTube, the same POV structure should carry a professional human voiceover instead, since a longer-attention audience reads TikTok-tier TTS as cheap rather than authentic. Testing both would show whether the platform, not the concept, should decide voice quality.
