So, I took a quick side quest from the book writing to record some music tracks! Here is ‘Sideways Again‘ – a song about bumping into a few roadblocks along the way, but keeping your head up because things always have a way of working out in the end.
Give it a listen and let me know what you think! 👇🎵
As an artist, I get excited about my creations, and I show my art to my close friends and family. And you know what I heard one of them say: “Oh, that’s AI slop. That’s a nice hobby.” WTF?
That stung and breaks my heart. Alas, recently heard “When it comes to humanity, disappointment is where you should start, not end up” Haha.. and I thought that was pretty good. Meaning: the only way is up.
And with that in mind, I thought I’d educate everyone on why this is not just merely an AI slop hobby:

How I Made “Sideways Again” the Music, Song, and Video

First, I made the music.

Phase 1: Pre-Production and Composition (The Human Blueprint)

Before sending a single prompt into Suno, the foundational elements of the song were entirely human-driven. Rather than relying on AI to generate it’s own ideas at random, the musical and textual DNA was locked down first:

a. Lyric Writing: Drafted the full lyric structure, rhythm, and rhyme schemes to control the song’s thematic arc and vocal pacing.

b. Harmonic and Melodic Framework: Established the key signature, exact chord progression, and melodic contours to define the song’s emotional identity.

c. Arrangement and Instrumentation: Decided the specific instrumental palette, section transitions (intro, verse, chorus, bridge, outro), and overall energy curves for the virtual band to follow.

Phase 2: Seed Recording and Directing the AI (The Hybrid Engine)

To ensure Suno didn’t just guess what the song should sound like, I put direct audio inputs and explicit text parameters in to guide the generation engine. As you can see from the pic above, there were many iterations before the ‘final rough cut.”

Here’s what I gave it:

a. Guide Track Recording: Recorded initial reference tracks of human guitar playing and vocal melodies to serve as the structural anchor.

b. Audio-Guided Generation: Fed those real audio recordings into Suno so the AI could analyze, track, and build instrumental arrangements directly around the actual feel and timing of the played instrument.

c. Prompt Direction: Specified the exact instrumentation, tone, and genre parameters, effectively acting as the producer instructing session musicians on what parts to play.

Phase 3: Deconstruction and Track Extraction to Transfer to DAW 

Once the best audio generated candidate passes based on my input guides, the output was treated not as a finished master, but as raw studio tracking:

a. Stem (Track) Separation: Downloaded the isolated stem tracks (drums, bass, secondary instruments, ambient textures) out of the music app.

b. DAW Import: Imported all individual multi-track stems into a Digital Audio Workstation (DAW) to regain full control over volume, panning, timing, and processing. I use Garage Band because I’m a Mac nerd and it is enough for my demos.

Phase 4: Live Human Overdubs and Performance

With the virtual band backing tracks aligned inside the DAW, true live performances were tracked to bring human dynamics and feel to the forefront:

a. Vocal Tracking: Sang and recorded the official lead vocal tracks direct to tape in the studio, capturing genuine human emotion, breath, and delivery that AI cannot replicate.

b. Guitar Overdubs: Laid down final, crisp human guitar tracks over the AI-generated backing to ground the rhythm section with organic feel and timbre.

Phase 5: Post-Production, Mixing, and Mastering (The Craft)

This is actually what I went to school for. I said to heck with University, I’m going down to Florida to enroll in the Audio Engineering and Recording program at Full Sail University. And I did. Learned a lot about this stuff.

The final polish of this song relied entirely on traditional audio engineering techniques across both individual channels and the master bus:

a. Noise Gate: Applied to individual stems and live tracks to clean up background noise, room bleed, and breath spills on live vocal and guitar recordings.

b. Pitch Correction: Used on individual vocal tracks to micro-tune live takes for precise intonation while maintaining natural human timbre.

c. Individual Track EQ (Equalization): Carved out mud, removed harsh resonances, and separated frequency ranges across vocal and instrumental stems so every instrument sat clearly in the mix without clashing.

d. Individual Track Compression: Tamed dynamic peaks on the live vocal and glue-compressed individual instrumental stems for punch and consistency.

e. Echo and Reverb: Applied to individual tracks to place the lead vocal and instruments into a realistic 3D spatial room.

f. Master Bus EQ: Applied across the combined mix to balance the overall final tonal curve, sweetening high-end air and solidifying low-end punch.

g. Master Bus Compression: Applied gently across the main stereo output to glue all human and AI elements into a single cohesive track.

h. Master Bus Limiter: Applied to the master channel to control peak levels and achieve commercial loudness standards without distortion or clipping.

Second, I made the video.

1. The Song & Initial Concept

It started with a Soul-Pop / Jazzy Soft Rock track called “Sideways Again” — a reflective, cinematic song about a man navigating a romantic near-miss in New York City. The lyrics carried a surreal, disoriented undertone captured in the recurring hook: “looks like I’m sideways again.” That line became the visual anchor for everything that followed.

The creative brief was clear from the start: New York City nights, checker taxi cabs, a long black wool coat, light trail effects, and anamorphic streak flares — the kind of neo-noir, Fincher-cool visual language that makes a city feel alive and slightly dreamlike.

2. Story Development — Four Chapters, Then a Remix

The initial narrative was structured as four non-linear chapters — a structural approach inspired by David Fincher’s cool wit. The story followed The Drifter, a tall, salt-and-pepper man in his 40s, as he moved through a series of emotionally charged NYC encounters: a chance meeting on a film set, a rooftop bar, a hopeful second date attempt at a Chinatown brownstone, and a surreal resolution.

After the first full production pass, the narrative was restructured into a chronological remix — a more emotionally legible arc running approximately 3:30:

  • Act I — The Meet-Cute (cab in rain, film set stop, PA encounter, rooftop bar)
  • Act II — The Anticipation & No-Show (mirror prep, the Katheryn note, bodega flowers, dark windows, payphone)
  • Act III — The Return (rooftop solo, reflection, closing)

The chapter title cards were ultimately removed to let the story breathe as a continuous film.

3. Characters & Visual References

Two characters were central:

  • The Drifter — the male lead, defined by a user-uploaded frontal photo reference. Medium-short salt-and-pepper hair swept back, neat goatee, tall frame. This reference photo was used as the anchor for all AI character generation throughout.
  • The Disruption — the female lead, similarly anchored to user-uploaded photos showing her natural warmth and casual elegance.

Character consistency across AI-generated video is one of the hardest challenges in this medium — maintaining the same face, hair, and posture across dozens of independently generated clips required constant referencing back to the original photos and iterative refinement.

4. Storyboard Planning — The Creative Canvas

Before a single video was generated, the entire film was mapped out in a structured storyboard canvas — 8 scenes, approximately 59 shots, ~149 seconds total. Each shot received:

  • A detailed audiovisual description — shot size, camera movement, lighting, performance direction, art department notes
  • A duration aligned to the music’s phrase structure
  • A scene preview sketch generated as a visual reference

This canvas served as the single source of truth throughout production — every generation decision traced back to it.

5. Visual Style — Neo-Noir NYC

The visual design language was locked in early:

  • Color palette: Deep blacks, warm amber streetlamp glow, cool blue-green ambient city light, wet pavement reflections
  • Texture: Anamorphic streak flares on strong accent beats, light trail effects on moving traffic, shallow depth of field with cinematic bokeh
  • Mood: Fincher-cool restraint — unhurried, deliberate framing, never showy
  • The Sideways Effect: A recurring surreal motif where reflections, the physical world, and skylines smoothly rotate 90° to horizontal — dreamlike, gravitational, cinematic. This effect was first realized in a NYC skyline spiral clip that became the canonical visual reference for all subsequent “goes sideways” moments throughout the film.

6. Video Generation — The Tools

All video clips were generated using AI video models accessed through the VidMuse platform:

  • Seedance 2.0 Mini — the primary workhorse for narrative/cinematic shots (cost-efficient, capable of strong image-to-video and text-to-video generation at 720p)
  • Seedance 2.0 Pro — used for higher-fidelity shots requiring stronger character consistency
  • Kling Pro / Hailuo H3 / OmniHuman — deployed selectively for lipsync performance shots, character-driven close-ups, and scenes requiring stronger motion quality
  • Seedream — used for image generation (reference images, storyboard sketches, character variants)

The generation workflow for each clip followed a consistent pattern:

  1. Compile a detailed prompt from the shot description, visual style, and character references
  2. Select the appropriate model based on the shot type and budget
  3. Generate, review, and iterate — often 2–4 takes per clip
  4. Write generated file paths back to the project canvas for timeline assembly

Over the course of production, 20+ individual video clips were generated, reviewed, and refined — covering apartment interiors, NYC street scenes, taxi interiors, rooftop bars, Chinatown alleys, bodega interiors, brownstone stoops, and surreal mirror/reflection moments.

7. The Lipsync Moments

Two deliberate lipsync performance beats were woven into the film — the Drifter singing the chorus “looks like I’m sideways again” directly to camera, dressed in Ray-Ban sunglasses against two distinct NYC backgrounds (a rooftop bokeh and a Chinatown street). These were generated using portrait-driven audio-synced video models with the character reference photos as anchors, then carefully placed at emotionally resonant points in the timeline.

 

. 

8. Post-Production — iMovie Assembly & the Sideways Effect

With all AI-generated clips exported from the canvas, the clips were transferred into iMovie for final editorial assembly. In iMovie:

  • Clips were sequenced and trimmed to align with the music’s phrase and beat structure
  • Still photographs were integrated alongside video clips for textural variety
  • The “Sideways Effect” — the film’s central visual motif — was created and applied at key emotional moments by rotating and tilting footage to match the lyrical theme of disorientation. This visual treatment gave the surreal rotations a handcrafted, intentional quality that complemented the AI-generated imagery.
  • The final audio mix layered the Suno track throughout, with all video clips playing muted to keep the music front and center.

9. Reflections on the Process

This project is a demonstration of human-AI creative collaboration at its most hands-on. Every creative decision — the story, the characters, the visual grammar, the emotional beats — came from a human director’s vision. The AI tools served as a production team: generating footage on demand, iterating on notes, and making the impossible (a full music video produced solo) achievable.

The limitations were real: character consistency across AI generations required constant vigilance, subtle details (hand proportions, door handles, seat orientations) needed multiple retakes, and the surreal effects that felt clearest in the imagination were the hardest to describe in prompts. But the results — a cohesive, atmospheric, ~3:30 narrative music video — speak to where AI-assisted filmmaking is today.

The Workflow Realized

This hybrid workflow turns AI into a customizable studio tool rather than an automated generator. By controlling the initial composition, using live instrument seeds, extracting multi-track stems, tracking real human vocals, and applying professional mix-bus engineering, the result isn’t an “AI song”—it’s a fully realized, human-directed production that leverages cutting-edge tools to bring an original creative vision to life.

FINAL COMMENTARY

The true shift today isn’t just in the tech—it’s in the economics of artistic freedom. Historically, executing a vision of this scale meant overcoming massive financial barriers: a producer would easily spend $450 per session musician just to get a band into the tracking room, followed by an estimated $20,000 trip to New York City to shoot a high-end music video. Today, that entire paradigm has been flipped.

While new tools provide the virtual backing band and video generators act as the visual production crew, the real creative weight remains entirely human.

Despite the technical hurdles—from fine-tuning stems in the DAW to coaxing precise visual details like character consistency and hands out of video prompts—the final result is a seamless, atmospheric 3:30 narrative where every sound and frame answers directly to a single director-ME.

By combining custom audio stems, live performances, meticulous mix engineering, and guided generative visuals, the independent creator now commands the output of a full-scale record label and production house from a single desk. The playing field hasn’t just been leveled; the power balance has permanently shifted:

The artist no longer needs the studio, it is the studios that need the artist.

This is Not My “Hobby”—It is My “Craft”

I have been playing guitar and writing songs since I was 12 years old, and now, in my 50s, that lifelong dedication remains unchanged. My journey through music wasn’t built on a weekend whim; I cut my teeth performing in numerous live bands throughout my 20s, spent decades recording original tracks, and directed my own videos along the way. To sharpen my technical foundation, I attended Full Sail University for audio engineering and spent years in the trenches running live sound for club acts, mastering the art of mixing when analog gear and physical tape were the only options. Digital technology didn’t create my passion—it evolved alongside it.

After a lifetime of honing my ears, hands, instincts, and skills, I am now thrilled to watch technology finally catch up with my art.

Chip Von Gunten

Exploring the known and the unknown with a beat writer’s eye for truth