Video Input
input
NodeTool Cloud is in alpha — run workflows in the browser, on your own keys.NodeTool Cloud is in alpha. Try it →
Pull the spoken words out of a video. Audio is extracted locally, then transcribed. Note the provider choice: fal and Replicate both list speech-to-text models in their manifests, but neither provider implements the transcription capability, so those entries cannot execute. OpenAI, Gemini, HuggingFace, MiniMax and Together do implement it.
Let us review the coffee subscription launch. Sam will finish the landing page by Friday. Alex will test checkout on Safari before Tuesday. We agreed to offer both 200 gram and 500 gram bags. Nine pilot orders arrived late, so tracking emails are the top priority. Our next meeting is Thursday at 10. Please bring the shipping update and the final product photos.
Install the free desktop app for macOS, Windows, or Linux. It runs on your own machine, no account required to start.
Browse the built-in template library inside Studio and open this workflow onto the canvas. Every node is already wired up.
Connect the providers this workflow uses (Extract Audio, Video Input). Bring your own keys — you pay the provider directly.
Hit Run to execute the graph and watch results stream in. Swap models, edit prompts, or rewire nodes to make it yours.
Explore a related recipe with guided setup and editable storyboards or scripts.









Free, open source, and yours to run. Download Studio, open the template, and run it with your own keys.