Music Video Visualizer
Turn any song into a mood-matched music video. Whisper transcribes the lyrics, an LLM creative director reads the emotional arc, writes one image prompt per frame, FLUX renders every frame, and they are stitched back to your original audio. Differentiator: a full transcribe → analyze → fan-out → render → reassemble media pipeline in a single graph — no chat box can do this. Cost note: generates one image per frame (default 8) plus a Whisper transcription, so a run costs several fal-ai calls.

The workflow
Nodes in this workflow
15 nodes · 12 types- ×3String Inputnodetool.input.StringInput
- ×2Promptnodetool.text.Prompt
- Add Audionodetool.video.AddAudio
- Agentnodetool.agents.Agent
- Audio Inputnodetool.input.AudioInput
- Automatic Speech Recognitionnodetool.text.AutomaticSpeechRecognition
- Collectnodetool.control.Collect
- For Eachnodetool.control.ForEach
- Frame To Videonodetool.video.FrameToVideo
- List Generatornodetool.generators.ListGenerator
- Outputnodetool.output.Output
- Text To Imagenodetool.image.TextToImage
How to run it
- 01
Download NodeTool Studio
Install the free desktop app for macOS, Windows, or Linux. It runs on your own machine, no account required to start.
- 02
Open the Music Video Visualizer template
Browse the built-in template library inside Studio and open this workflow onto the canvas. Every node is already wired up.
- 03
Add your keys
Connect the providers this workflow uses (Add Audio, Agent, Audio Input). Bring your own keys — you pay the provider directly.
- 04
Run and remix
Hit Run to execute the graph and watch results stream in. Swap models, edit prompts, or rewire nodes to make it yours.
Run Music Video Visualizer on your machine
Free, open source, and yours to run. Download Studio, open the template, and run it with your own keys.








