All templates
Template·Audio & Music

Audio To Image

Speak an image into existence: no keyboard needed. Whisper transcribes your audio, then FLUX renders the description as an image — the whole pipeline runs from a single voice note.

Audio To Image — example output from the NodeTool workflow

The workflow

Workflow Editor
Note
Audio Input
Automatic Speech Recognition
Audio
Text
openai/whisper-large-v3
Text To Image
Prompt
fal-ai/flux/schnell
Output

Nodes in this workflow

4 nodes · 4 types
  • Audio Input
    nodetool.input.AudioInput
  • Automatic Speech Recognition
    nodetool.text.AutomaticSpeechRecognition
  • Output
    nodetool.output.Output
  • Text To Image
    nodetool.image.TextToImage

How to run it

  1. 01

    Download NodeTool Studio

    Install the free desktop app for macOS, Windows, or Linux. It runs on your own machine, no account required to start.

  2. 02

    Open the Audio To Image template

    Browse the built-in template library inside Studio and open this workflow onto the canvas. Every node is already wired up.

  3. 03

    Add your keys

    Connect the providers this workflow uses (Audio Input, Text To Image). Bring your own keys — you pay the provider directly.

  4. 04

    Run and remix

    Hit Run to execute the graph and watch results stream in. Swap models, edit prompts, or rewire nodes to make it yours.

Run Audio To Image on your machine

Free, open source, and yours to run. Download Studio, open the template, and run it with your own keys.