NodeTool Cloud is in alpha — run workflows in the browser, on your own keys.NodeTool Cloud is in alpha. Try it →
Transcribe a Clip
Pull the spoken words out of a video. Audio is extracted locally, then transcribed. Note the provider choice: fal and Replicate both list speech-to-text models in their manifests, but neither provider implements the transcription capability, so those entries cannot execute. OpenAI, Gemini, HuggingFace, MiniMax and Together do implement it.

The workflow
Nodes in this workflow
4 nodes · 4 types- Automatic Speech Recognitionnodetool.text.AutomaticSpeechRecognition
- Extract Audionodetool.video.ExtractAudio
- Outputnodetool.output.Output
- Video Inputnodetool.input.VideoInput
How to run it
- 01
Download NodeTool Studio
Install the free desktop app for macOS, Windows, or Linux. It runs on your own machine, no account required to start.
- 02
Open the Transcribe a Clip template
Browse the built-in template library inside Studio and open this workflow onto the canvas. Every node is already wired up.
- 03
Add your keys
Connect the providers this workflow uses (Extract Audio, Video Input). Bring your own keys — you pay the provider directly.
- 04
Run and remix
Hit Run to execute the graph and watch results stream in. Swap models, edit prompts, or rewire nodes to make it yours.
Part of a bigger job
A recipe chains this workflow with the ones around it and ships the whole set as a single file.
Related templates









Run Transcribe a Clip on your machine
Free, open source, and yours to run. Download Studio, open the template, and run it with your own keys.