Replicate logoModel provider · Bring your own key

Replicate in a visual AI workflow

Run Replicate's image, video, and audio models in a visual AI workflow — hundreds of community and first-party models, run with your own key at Replicate's list price.

ImageVideoAudioreplicate.com
655
Models
310
Image
156
Video
62
Audio
127
Text

Replicate hosts a large, fast-moving catalog of open and commercial models behind one API — image, video, and audio generators from FLUX and Stable Diffusion to Kling, Wan, and more. NodeTool surfaces each as a node you can drop onto the canvas.

Chain Replicate models into a pipeline — prompt to image, image to video, transcript to speech — and the whole graph is reusable and shareable. Because every model is a node, comparing two Replicate models on the same prompt is a matter of wiring both into the same input.

Replicate runs on your own key in NodeTool: set your `REPLICATE_API_TOKEN` and calls go straight to Replicate at their price. The model list below comes from the Replicate node manifest, so it tracks what the provider ships.

Set your key in NodeTool
REPLICATE_API_TOKEN

Requests go straight to Replicate with your key at their list price. NodeTool never sits in the middle and never adds a markup.

Supported Replicate models

Each model below is its own node in NodeTool — the id is what Replicate serves it under. This list is generated from Replicate's node manifest, so it tracks what the provider actually ships.

Image models

310 total
Background Remover 851
851-labs/background-remover

Remove backgrounds from images.

Mirage Ghibli
aaronaftab/mirage-ghibli

Ghiblify any image, 10x cheaper/faster than GPT 4o

Flux Cinestill
adirik/flux-cinestill

Flux lora, use "CNSTLL" to trigger

Kandinsky
ai-forever/kandinsky-2.2

multilingual text2image latent diffusion model

Alexgenovese Upscaler
alexgenovese/upscaler

GFPGAN aims at developing Practical Algorithms for Real-world Face and Object Restoration

Ra2
anon987654321/ra2
Deoldify Image
arielreplicate/deoldify_image

Add colours to old images

Dwiss Qwen 2
arthuryeti/dwiss-qwen-2
Flux 1 1 Pro Ultra
black-forest-labs/flux-1.1-pro-ultra

FLUX1.1 [pro] in ultra and raw modes.

Flux 1 1 Pro Ultra Finetuned
black-forest-labs/flux-1.1-pro-ultra-finetuned

Inference model for FLUX 1.1 [pro] Ultra using custom `finetune_id`.

Flux 2 Dev
black-forest-labs/flux-2-dev

Quality image generation and editing with support for reference images

Flux 2 Flex
black-forest-labs/flux-2-flex

Max-quality image generation and editing with support for ten reference images

Flux 2 Klein 4 B
black-forest-labs/flux-2-klein-4b

Very fast image generation and editing model.

Flux 2 Klein 4b Base
black-forest-labs/flux-2-klein-4b-base

Un-distilled version of FLUX.2 [klein].

Flux 2 Klein 4b Base Lora
black-forest-labs/flux-2-klein-4b-base-lora

A version of FLUX.2 [klein] 4B-base that supports fast fine-tuned lora inference

Flux 2 Klein 9b
black-forest-labs/flux-2-klein-9b

4 step distilled version of FLUX.2 [klein].

Flux 2 Klein 9b Base
black-forest-labs/flux-2-klein-9b-base

Un-distilled version of FLUX.2 [klein].

Flux 2 Klein 9b Base Lora
black-forest-labs/flux-2-klein-9b-base-lora

A version of FLUX.2 [klein] 9B-base that supports fast fine-tuned lora inference

Flux 2 Max
black-forest-labs/flux-2-max

The highest fidelity image model from Black Forest Labs

Flux 2 Pro
black-forest-labs/flux-2-pro

High-quality image generation and editing with support for eight reference images

Flux Canny Dev
black-forest-labs/flux-canny-dev

Open-weight edge-guided image generation.

Flux Canny Pro
black-forest-labs/flux-canny-pro

Professional edge-guided image generation.

Flux Depth Dev
black-forest-labs/flux-depth-dev

Open-weight depth-aware image generation.

Flux Depth Pro
black-forest-labs/flux-depth-pro

Professional depth-aware image generation.

Flux Dev
black-forest-labs/flux-dev

A 12 billion parameter rectified flow transformer capable of generating images from text descriptions

Flux Dev Lora
black-forest-labs/flux-dev-lora

A version of flux-dev, a text to image model, that supports fast fine-tuned lora inference

Flux Fill Dev
black-forest-labs/flux-fill-dev

Open-weight inpainting model for editing and extending images.

Flux Fill Pro
black-forest-labs/flux-fill-pro

Professional inpainting and outpainting model with state-of-the-art performance.

Flux Kontext Dev
black-forest-labs/flux-kontext-dev

Open-weight version of FLUX.1 Kontext

Flux Kontext Dev Lora
black-forest-labs/flux-kontext-dev-lora

FLUX.1 Kontext[dev] image editing model for running lora finetunes

Flux Kontext Max
black-forest-labs/flux-kontext-max

A premium text-based image editing model that delivers maximum performance and improved typography generation for transforming images through natural languag…

Flux Kontext Pro
black-forest-labs/flux-kontext-pro

A state-of-the-art text-based image editing model that delivers high-quality outputs with excellent prompt following and consistent results for transforming…

Flux Krea Dev
black-forest-labs/flux-krea-dev

An opinionated text-to-image model from Black Forest Labs in collaboration with Krea that excels in photorealism.

Flux Pro
black-forest-labs/flux-pro

State-of-the-art image generation with top of the line prompt following, visual quality, image detail and output diversity.

Flux Pro Finetuned
black-forest-labs/flux-pro-finetuned

Inference model for FLUX.1 [pro] using custom `finetune_id`

Flux Redux Dev
black-forest-labs/flux-redux-dev

Open-weight image variation model.

Flux Redux Schnell
black-forest-labs/flux-redux-schnell

Fast, efficient image variation model for rapid iteration and experimentation.

Flux Schnell
black-forest-labs/flux-schnell

The fastest image generation model tailored for local development and personal use

Flux Schnell Lora
black-forest-labs/flux-schnell-lora

The fastest image generation model tailored for fine-tuned use

Bria Eraser
bria/eraser

SOTA Object removal, enables precise removal of unwanted objects from images while maintaining high-quality outputs.

+ 270 more image models available in Replicate.

Video models

156 total
Happy Horse 1
alibaba/happyhorse-1.0

Alibaba's Happy Horse 1.0 generates videos from text prompts or animates a single image into video.

Happyhorse 1 1
alibaba/happyhorse-1.1

Alibaba's Happy Horse 1.1 generates videos from text, animates a single image, or builds a video from multiple reference images.

Wan 1 3b Inpaint
andreasjansson/wan-1.3b-inpaint

Inpainting and video2video experiments with Wan 2.1

Zeroscope V2 XL
anotherjesse/zeroscope-v2-xl

Zeroscope V2 XL & 576w

Deoldify Video
arielreplicate/deoldify_video

Add colours to old video footage.

Robust Video Matting
arielreplicate/robust_video_matting

extract foreground of a video

Video Erase Object
bria/video-erase-object

A high-fidelity capability for erasing unwanted objects, people, or visual elements from videos while maintaining aesthetic quality and temporal consistency

Video Increase Resolution
bria/video-increase-resolution

Upscale videos up to 8K output resolution.

Video Remove Background
bria/video-remove-background

Automatically remove backgrounds from videos -perfect for creating clean, professional content without a green screen.

Dream Actor M2
bytedance/dreamactor-m2.0

Animate any character, humans, cartoons, animals, even non-humans, from a single image + driving video

Latent Sync
bytedance/latentsync

LatentSync: generate high-quality lip sync animations

Omni Human
bytedance/omni-human

Turns your audio/video/images into professional-quality animated videos

Omni Human 1 5
bytedance/omni-human-1.5

A film-grade digital human model that generates realistic video from a single image, audio clip, and optional text prompt.

Seedance 1 Lite
bytedance/seedance-1-lite

A video generation model that offers text-to-video and image-to-video support for 5s or 10s videos, at 480p and 720p resolution

Seedance 1 Pro
bytedance/seedance-1-pro

A pro version of Seedance that offers text-to-video and image-to-video support for 5s or 10s videos, at 480p and 1080p resolution

Seedance 1 Pro Fast
bytedance/seedance-1-pro-fast

A faster and cheaper version of Seedance 1 Pro

Seedance 1 5 Pro
bytedance/seedance-1.5-pro

A joint audio-video model that accurately follows complex instructions.

Seedance 2
bytedance/seedance-2.0

ByteDance's multimodal video generation model with native audio, multimodal reference inputs, and intelligent duration control.

Seedance 2 0 Fast
bytedance/seedance-2.0-fast

A faster variant of Seedance 2.0 for quicker video generation with multimodal inputs and native audio.

Seedance 2 0 Mini
bytedance/seedance-2.0-mini

A lower-cost variant of Seedance 2.0 for high-volume video generation with multimodal inputs and native audio.

Video Upscaler
bytedance/video-upscaler

Upscale and enhance video up to 4K at 60fps, with scene-aware presets for AI-generated content, short dramas, UGC, and film restoration.

Ovi I2v
character-ai/ovi-i2v

Ovi: generate videos with audio from image and text inputs

Video Retalking
chenxwh/video-retalking

Audio-based Lip Synchronization for Talking Head Video

Ani Portrait
cjwbw/aniportrait-audio2vid

Audio-Driven Synthesis of Photorealistic Portrait Animations

Sad Talker
cjwbw/sadtalker

Stylized Audio-Driven Single Image Talking Face Animation

Lucy Edit 2
decart/lucy-edit-2

Edit and transform videos with text prompts and reference images.

Ai Avatars
easel/ai-avatars

Use one or two face images to create AI avatars

Auto Caption
fictions-ai/autocaption

Automatically add captions to a video

Restyle Video Frame
flux-kontext-apps/restyle-video-frame

Use flux-kontext-pro to change the first or last frame of a video.

Audio To Waveform
fofr/audio-to-waveform

Create a waveform video from audio

Kontext Ps1
fofr/kontext-ps1

FLUX Kontext fine-tune that let's you restyle any image as a PS1 or PS2 video game

Veo 2
google/veo-2

State of the art video generation model.

Veo 3
google/veo-3

Sound on: Google’s flagship Veo 3 text to video model, with audio

Veo 3 Fast
google/veo-3-fast

A faster and cheaper version of Google’s Veo 3 video model, with audio

Veo 3 1
google/veo-3.1

New and improved version of Veo 3, with higher-fidelity video, context-aware audio, reference image and last frame support

Veo 3 1 Fast
google/veo-3.1-fast

New and improved version of Veo 3 Fast, with higher-fidelity video, context-aware audio and last frame support

Veo 3 1 Lite
google/veo-3.1-lite

Google's cost-efficient video generation model with native audio, optimized for high-volume applications

Avatar Iv
heygen/avatar-iv

Create realistic talking avatar videos from text with HeyGen's Avatar IV engine

Avatar V
heygen/avatar-v

Create realistic talking avatar videos from text with HeyGen's Avatar V engine — the newest, highest-quality avatar engine with cross-reference-driven animat…

Hey Gen Lipsync Precision
heygen/lipsync-precision

High-accuracy lip-sync: replace or dub audio on any video with avatar-inference lip sync

+ 116 more video models available in Replicate.

Audio models

62 total
Style TTS2
adirik/styletts2

Generates speech from text

Tortoise TTS
afiaka87/tortoise-tts

Generate speech from text, clone voices from mp3 files.

Music Gen Looper
andreasjansson/musicgen-looper

Generate fixed-bpm loops from text prompts

Open Voice
chenxwh/openvoice

Updated to OpenVoice v2: Versatile Instant Voice Cloning

Parler TTS
cjwbw/parler-tts

lightweight text-to-speech (TTS) model, trained on 10.5K hours of audio data

Voice Craft
cjwbw/voicecraft

Zero-Shot Speech Editing and Text-to-Speech in the Wild

Eleven Labs Flash V2 5
elevenlabs/flash-v2.5

ElevenLabs's fastest speech synthesis model

Eleven Labs Music
elevenlabs/music

Compose a song from a prompt or a composition plan

Eleven Labs Turbo V2 5
elevenlabs/turbo-v2.5

High quality, low latency text to speech in 32 languages

Eleven Labs V2 Multilingual
elevenlabs/v2-multilingual

Generate multilingual text-to-speech audio in over 30 languages

Eleven Labs V3
elevenlabs/v3

The most expressive Text to Speech model

Spanish F5 TTS
fermatresearch/spanish-f5-tts

A F5-TTS fine-tuned for Spanish

Gemini 3 1 Flash TTS
google/gemini-3.1-flash-tts

Google's fast, expressive text-to-speech model with 30 voices and 70+ language support

Lyria 2
google/lyria-2

Lyria 2 is a music generation model that produces 48kHz stereo audio through text-based prompts

Lyria 3
google/lyria-3

Generate 30-second music clips from text prompts or images with Lyria 3, Google's music generation model

Lyria 3 Pro
google/lyria-3-pro

Generate full-length songs up to 3 minutes from text prompts or images with Lyria 3 Pro, Google's most capable music generation model

Realtime Tts 1 5 Max
inworld/realtime-tts-1.5-max

Highest-quality realtime text-to-speech with <200ms latency, emotion control, and 15-language support

Realtime Tts 1 5 Mini
inworld/realtime-tts-1.5-mini

Ultra-fast, cost-efficient realtime text-to-speech with ~120ms latency and 15-language support

Inworld Realtime TTS 2
inworld/realtime-tts-2

Most expressive text-to-speech model from Inworld, with natural-language steering, real-time latency, and multilingual support across 100+ languages.

Inworld TTS Max
inworld/tts-1.5-max

Highest-quality text-to-speech with <200ms latency, emotion control, and 15-language support

Inworld TTS Mini
inworld/tts-1.5-mini

Ultra-fast, cost-efficient text-to-speech with ~120ms latency and 15-language support

Kokoro 82 M
jaaari/kokoro-82m

Kokoro v1.0 - text-to-speech (82M params, based on StyleTTS2)

Ace Step
lucataco/ace-step

A Step Towards Music Generation Foundation Model text2music

CSM 1 B
lucataco/csm-1b

CSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs

+ 38 more audio models available in Replicate.

Text models

127 total
Text Extract OCR
abiruyt/text-extract-ocr

A simple OCR Model that can easily extract text from an image.

E5 Mistral 7 B
adirik/e5-mistral-7b-instruct

E5-mistral-7b-instruct language embedding model

Blip2
andreasjansson/blip-2

Answers questions about images

Clip Features
andreasjansson/clip-features

Return CLIP features for the clip-vit-large-patch14 model

CLIP Features
andreasjansson/clip-features

Return CLIP features for the clip-vit-large-patch14 model

Llama2 13 B Embeddings
andreasjansson/llama-2-13b-embeddings

Llama2 13B with embedding output

Claude 3 5 Haiku
anthropic/claude-3.5-haiku

Anthropic's fastest, most cost-effective model, with a 200K token context window (claude-3-5-haiku-20241022)

Claude 3 7 Sonnet
anthropic/claude-3.7-sonnet

The most intelligent Claude model and the first hybrid reasoning model on the market (claude-3-7-sonnet-20250219)

+ 119 more text models available in Replicate.

Frequently asked questions

How do I connect Replicate to NodeTool?
Add your Replicate token as `REPLICATE_API_TOKEN` in NodeTool. Replicate nodes then call the API directly with your token.
Which Replicate models are available?
The image, video, and audio models in Replicate's node manifest — several hundred, each a separate node. See the catalog below.
What does Replicate cost through NodeTool?
Replicate's own list price. NodeTool is bring-your-own-key and adds no per-generation fee.

Other providers

Run Replicate your way.

Download NodeTool Studio and build across image, video, audio, and text with your own keys.