On this page

Two ways to think#

Daniel Kahneman’s Thinking, Fast and Slow splits the mind in two. System 1 is fast, automatic, and confidently wrong about the bat and the ball. System 2 is slow, deliberate, and expensive.

For most of their short history, language models answered like System 1: one forward pass per token, and the first thought is the answer. Then in 2024, OpenAI’s o1 showed that a model keeps getting better the longer it thinks before answering. Thinking time became a second way to scale AI, right next to training.

Nice & Smooth’s 1991 classic “Sometimes I Rhyme Slow” was already that duo. Greg Nice is quick and bouncy; Smooth B is low, silky, and takes his time. Swap one word, and the hook explains the whole idea.

Watch original video
The original: Nice & Smooth, “Sometimes I Rhyme Slow” (1991)

It’s a follow-up to Slow It Down, my first AI-made music video, and it was made the same way: Claude Code directing, me giving notes.

Meet FAST & SLOW#

“Kahneman wrote a whole book on why I’m problematic.”

FAST is System 1: a skater with a wild mane who answers before you finish asking. SLOW is System 2: a ronin with a notebook who takes three minutes and gets it right. They’re twins (same data, same weights) with the love-hate energy of Samurai Champloo’s Mugen and Jin: they bicker, roast each other, and work best together.

Verse 1 is FAST’s world, in saturated 90s color: two thousand tokens a second, “low effort still beat your high,” Andrej Karpathy’s “no thinking, single token, low latency” on a pager, a confident “twenty twenty-nine” for today’s date, and the bat and the ball. (Ball’s a dime. It’s a nickel.)

“Low effort still beat your high,” and a page from Karpathy.

Verse 2 is SLOW’s story, in black-and-white sumi-e ink wash, with FAST as the only color in his memories. SLOW was ten thousand agents deep on Navier–Stokes while his twin was in the replies cracking jokes, telling the whole world to put glue on pizza, and special-casing the tests. So SLOW sends him off to RL with verifiable goals, and a year later he comes home… guessing again.

SLOW, ten thousand agents deep on Navier–Stokes. FAST, on a banana peel under a Gary Marcus headline.

By the last hook it’s sunrise, and the twins share one mixing desk: one fader on LOW, one on MAX.

Real people only ever appear as in-world text: a pager, a CRT, a newspaper headline, a book cover. The captions live in the world too. FAST’s lines are spray-painted tags, SLOW’s are handwriting, and the hook is painted sign lettering.

How it was made#

Claude Code with Opus 5.5 made nearly all of it: the research, the lyrics, the art direction, the storyboard, every image and video prompt, the compositor, and the review tools. My job was taste and feedback, at three checkpoints: the art direction, the storyboard, and the video. Start to finish took about a day of back-and-forth and roughly $60 of fal credits.

  1. Lyrics. Rewritten line by line against the original’s bars, syllables, and rhyme scheme, and packed with references from the last two years of AI.
  2. Song. Generated from scratch in Suno V6 from text alone: our lyrics plus a style prompt describing a laid-back 1991 New York boom-bap record, without naming the original.
  3. Timing. Word-level timestamps for the Suno take, so every caption lands on its word.
  4. Characters. One model sheet each for FAST and SLOW, passed into every image.
    The first model sheets (top), and the final ones (bottom) after I asked for more Mugen and Jin Samurai Champloo vibes.
  5. Storyboard. 66 shots, each starting on its lyric, with one keyframe per shot from Nano Banana Pro. The color follows the mode: dusk for the hooks, saturated for FAST, ink wash for SLOW, and gold at sunrise.
    The approved storyboard: 66 keyframes in song order, color-coded by mode.
  6. Clips. Kling 3 Pro for color and action, and Wan 3.0 for SLOW’s calmer ink-wash shots.
  7. Compositing. A local compositor lays the clips on the song, grades each mode, adds VHS grain and scanlines, and animates every caption word by word.
  8. Review. A scene-by-scene review tool with Approve and Change on every shot, plus a notes box. 59 of 66 shots passed the first full render, and my notes on the other seven became round two.
    Round two: my real note on the crew shot, and what changed. Show render 1 swaps back to the old take.
  9. Poster. The end card doubles as the first frame, since X shows a video’s first frame as its preview.

Tried and cut#

None of this made it into the final video.

  • Voice auditions. Before Suno, I auditioned AI rap voices from Lyria 3.5 and 3 Pro, MiniMax Music 2.6 and 3, ACE-Step, ElevenLabs, and Seed-VC. None of them had the laid-back 1991 delivery.
  • Lip-sync. OmniHuman 1.5 lip-synced nine close-ups. The cut plays better without them.
  • Other video models. Veo 3.1 Lite was faithful but tamer, and PixVerse v6 reframed shots and drifted off-model.
  • Other image models. Seedream 5 Pro was in the first style bake-off, next to Krea 2 and Nano Banana Pro.

What I learned#

  • Video models rewrite in-world text. Kling kept “improving” FAST’s 2029 until it read 4029. Pinning the approved keyframe as the clip’s last frame fixed it, and anything that must never change gets pasted back from the keyframe.
    Same shot, same moment: render one’s 4029 (left), and the fix with the keyframe pinned as the last frame (right).
  • Watch the last second of every clip. That’s where faces drift off-model and cameras cut away. We used only the clean part, slowed to fit the shot.
  • One model per look. Kling was best for action and keeping faces on-model. Wan was just as good, at about a third of the price, for calm ink-wash shots.
  • Put real people in the world, not on screen. A pager, a CRT post, a newspaper, and a book cover carry every reference without anyone’s face.

Suno: what worked#

Suno made the song from text alone. For anyone making their own:

  • Spell names the way they should be sung. The lyrics say “Ahn-dray,” “Kah-nuh-mun,” and “Nav-yay Stokes,” and the captions switch back to the real spellings.
  • Describe the sound, never the artist. The style prompt gives the era, tempo, key, instruments, and delivery.
  • Spell out the hook in the style prompt too, call and response included.
The style prompt
plain text
1991 New York golden-era hip-hop, laid-back and sunny, boom-bap at 108 BPM in B major. Dusty crisp drum break with a fat snare on 2 and 4, warm round bass, and a mellow looped clean electric guitar arpeggio with a fingerpicked folk-pop feel. Male MC with a smooth, mellow, warm mid-low voice; relaxed, conversational, half-sung rap that sits slightly behind the beat. The hook is a chanted gang-vocal call-and-response with two voices and a punchy echo. The hook: the MCs chant "sometimes I think slow, sometimes I think fast" and the crew answers "fast, fast, fast". Dry, upfront, near-mono vocals, vinyl warmth, sampler grit, early-90s radio mix.
Exclude styles
plain text
trap, 808 slides, trap hi-hats, hi-hat rolls, ticking, clock, metronome, double-time, autotune, drill, EDM, dubstep, rock, metal, pop punk, orchestral, screaming, heavy reverb, female lead vocal

Parody. Not affiliated with Nice & Smooth, any AI lab, or anyone it name-checks.