Our AI Models
From concept art to full cinematic video, Layer brings together the industry's leading AI models in one place.
GPT Image 2
OpenAI's latest image model with stronger text rendering, UI generation, and photorealism. Native output up to 4K with three quality tiers.
Gemini Omni Flash
Google's fast multimodal video model — 720p clips with synchronized native audio from text or a still image.
Gemini Omni Flash Reference
Reference-guided video generation with up to 5 reference images for consistent characters and style.
MiniMax H3
MiniMax H3 (Hailuo-03) next-gen open-weight video with native stereo audio at 2K and 24 FPS.
MiniMax H3 Reference
Reference-guided MiniMax H3 video with multimodal image, video, and audio inputs at 2K.
MiniMax H3 Max
MiniMax H3 Max post-trained video with native stereo audio at 480p, 768p, 2K, or 4K.
Seedance 2.5 Reference
Reference-guided Seedance 2.5 video with up to 30 images, 10 videos, and 10 audio files.
Seedance 2
Professional-grade video model with cinematic quality, up to 4K output, and synchronized audio generation.
Seedance 2 Reference
Generate videos guided by reference images, videos, and audio with precise style and character control.
Seedance 2.5
ByteDance's Seedance 2.5 video model with up to 30s clips, 480p/720p/1080p output, and native audio.
Seedance 2 Fast
Fast, cost-effective variant of Seedance 2 with 720p output and synchronized audio generation.
Seedance 2 Fast Reference
Fast reference-guided video generation with multi-modal inputs and synchronized audio.
Grok Imagine Video 1.5
xAI's Grok Imagine 1.5 image-to-video model, animating images into 480p or 720p clips.
Grok Imagine Video
xAI's video generation model capable of creating high-quality 720p video from text and images.
Grok Imagine Video Extend
Extend an existing video with xAI's Grok Imagine, continuing motion from the source ending.
PixVerse v6
Latest PixVerse video model with improved quality and audio generation, supporting up to 1080p resolution.
Reve 2.1 Remix
Reve's remix model - blends one to eight reference images into a single new image, guided by a prompt.
Reve 2.1
Reve's next-generation image model with strong prompt adherence and text rendering, plus native editing.
Gemini 3.1 Flash Image
Google's fast, high-quality image generation model with multimodal reasoning.
GPT Image 1.5
GPT Image 1.5 is the latest image generation model from OpenAI, with better instruction following and adherence to prompts.
Happy Horse 1.1
Alibaba's text-to-video and image-to-video model generating 720p or 1080p clips with native audio and multilingual lip-sync.
MAI Image 2.5
Microsoft's photorealistic image model with strong typography and native editing.
Gemini 3 Pro Image
Google's state-of-the-art image generation and editing model.
MAI Image 2.5 Pro
Microsoft's flagship image model for high-fidelity generation and native editing.
Happy Horse 1.0
Alibaba's text-to-video and image-to-video model generating 720p or 1080p clips with native audio and multilingual lip-sync.
Happy Horse 1.0 Reference
Reference-guided video generation with up to 9 reference images for consistent characters and style.
Gemini 3.1 Flash-Lite Image
Google's fastest, most cost-efficient Gemini image generation model.
Kling v2.5 Turbo Pro
High-speed video model. Produces rapid, professional-quality 1080p video for fast-paced content.
Kling V3 Pro
Generate high-quality videos with advanced control. Pro tier with negative prompts.
Kling V3 4K
Generate native 4K videos directly from prompts or images. Cinema-grade output in one step.
Vidu Q3 Pro
Vidu's Q3 Pro video model with up to 16s clips, 720p/1080p output, and native audio.
Kling O3 Pro
Generate high-quality videos with prompts, images, or reference elements. Pro tier for premium quality.
Kling O3 4K Reference
Generate native 4K videos guided by reference images and elements. Cinema-grade output.
Kling O3 4K
Generate native 4K videos with Kling's Omni 3 model. Cinema-grade output in one step.
PixVerse v5.5
Latest PixVerse video model with enhanced quality and audio generation, supporting up to 1080p resolution.
PixVerse v5.5 Fast
Speed-optimized PixVerse v5.5 variant for rapid video generation at up to 720p resolution.
Kling O3 Standard
Generate videos with prompts, images, or reference elements. Standard tier for balanced quality and speed.
Kling V3 Standard
Generate videos with prompts and images. Standard tier with negative prompts.
PixVerse v5
Powerful, user-friendly video model producing high-quality, stylized 1080p video with consistent motion.
Kling v2.6 Pro
Generate videos from images with native audio generation and fluid motion.
Runway Gen-4.5
Runway's high-fidelity image-to-video model with improved prompt adherence and motion quality.
Veo 3.1 Fast
Speed-optimized Veo 3.1 model. Delivers high-quality video rapidly for dynamic workflows.
Veo 3.1 Fast Reference to Video
Speed-optimized reference-guided Veo 3.1: character-consistent video from up to three reference images.
Veo 3.1
Flagship video update: refined control, enhanced visual fidelity, and improved subtle motion details.
Veo 3.1 Reference to Video
Reference-guided Veo 3.1: lock characters and objects across shots using up to three reference images.
Seedance 1.5 Pro
Next-generation Seedance video model with improved quality, 1080p output, and integrated audio generation.
Veo 3.1 Lite
Cost-effective Veo 3.1 variant. Balances quality and affordability for high-volume video creation.
Minimax Hailuo-2.3 Pro
Pinnacle of Hailuo T2V series. Cinematic quality with superior coherence, detail, and artistic control.
Minimax Hailuo-2.3 Standard
Latest standard video model. Improved prompt understanding and visual consistency for daily creation.
Minimax Hailuo-02 Pro
Premium video model engineered for professional-grade 1080p output, superior fidelity, and smoother motion.
Qwen 3 TTS
Multilingual TTS with zero-shot voice cloning and prompt-based voice design.
Grok Imagine Image
xAI's image generation model capable of creating high-quality images from text prompts.
Qwen Image 2 Pro
Qwen Image 2 Pro tier with higher quality text-to-image generation and enhanced detail.
Wan 2.5
Powerful video model. Optimized for top-tier 1080p cinematic quality and consistency.
Minimax Hailuo-02 Standard
Robust video model producing crisp 768p video. Offers a solid balance of quality and reliable performance.
Seedance Pro
High-quality video model. Tuned for maximum visual fidelity and broadcast-quality 1080p output.
FLUX.2 [max]
Seedream 4.0
Unified architecture for image generation and editing. Allows fluid movement from concept to refinement.
FLUX.2 [flex]
Veo 3
Next-gen video model with enhanced control over narrative, tone, and shot composition. Includes audio.
Luma UNI-1 Max
The highest-fidelity tier of Luma UNI-1, for hero-quality stills and premium edits.
Ideogram V4
Ideogram's latest text-to-image model with crisp visuals, accurate text, and image-to-image support.
Krea 2 Large
Krea's flagship text-to-image model for high-fidelity generations with distinctive aesthetic range.
Krea 2 Turbo
Speed-optimized open-source version of Krea 2 — high-fidelity images in seconds, with the full Krea aesthetic range.
LTX Video 2.5 Fast
Speed-optimized LTX 2.5 audio-video model with 720p–4K output and clips up to 20 seconds.
LTX Video 2.5 Fast Audio-to-Video
Speed-optimized LTX 2.5 mode that generates video timed to a supplied audio clip.
Wan 2.6
State-of-the-art multimodal video generation model from Alibaba, with native audio support
Kling Image V3
Kling V3 image model. High-quality images with negative prompts, supports up to 2K resolution.
Gemini 3.1 Flash TTS
Google's most controllable TTS model with 200+ audio tags for vocal style and delivery.
FLUX.2 [pro]
Seedream 4.5
A new-generation image creation model from ByteDance, for both generation and editing.
LTX Video 2.5 Pro
Quality-optimized LTX 2.5 audio-video model for high-fidelity 720p and 1080p output.
LTX Video 2.5 Pro Audio-to-Video
Quality-optimized LTX 2.5 mode for final visuals synchronized to music, dialogue, or soundtrack.
GPT Image 1
Powerful, versatile OpenAI image model for creative and professional apps.
Wan 2.7 Pro
Alibaba WAN 2.7 Pro image model for high-quality text-to-image generation and multi-image editing.
Kling v2.1 Master
Premium 1080p video model. Maximum visual fidelity, capturing intricate details and lifelike expressions.
Luma UNI-1
Luma's unified image model for high-fidelity generation and prompt-based editing.
FLUX.2 [dev]
FLUX.2 [dev] Edit
Kling Image O3
Kling Omni 3 image model. High-quality images with text rendering capabilities up to 4K resolution.
Recraft V4.1
Professional design model with extended prompt support (10,000 characters) and advanced style customization.
Seedream 5.0 Lite
Fast, high-quality image generation from ByteDance, optimized for creative advertising.
Kling O1
Generate new videos from first and last frame images.
Kling O1 Reference
Generate new videos guided by prompts, images or videos.
Pika 2.5
Pika's highest-quality image-to-video model, generating 480p, 720p, or 1080p clips from a still frame.
Minimax Hailuo-2.3 Fast
High-speed T2V variant. Optimized for rapid creation, iteration, and social media content workflows.
HiDream O1 Image 1.0
HiDream's full O1 Image model — create, edit, and personalize images up to 2K in one native model.
Gemini 2.5 Flash Image
Google's image generation and editing model capable of multimodal reasoning.
LTX Video 2.0 Pro
High-fidelity video model designed for professional-quality results, offering superior detail and nuanced audio.
Recraft V4.1 Pro
Professional design model with extended prompt support (10,000 characters) and advanced style customization.
LTX Video 2.0 Fast
Versatile video model integrating video and audio creation in one seamless, speed-optimized workflow.
Kling v2.1 Pro
Refined Kling model delivering professional-grade 1080p video with improved clarity and motion.
Recraft V4
Professional design model with extended prompt support (10,000 characters) and advanced style customization.
Recraft V4 Pro
Recraft's pro V4 tier, with higher fidelity and finer detail than the standard tier.
ElevenLabs TTS V3
ElevenLabs' most expressive TTS model with inline audio tags for emotion and delivery.
Kling v2.0 Master
Master-grade video model. Enhanced realism and physics simulation for cinematic, high-impact clips.
Veo 3 Fast
Speed-optimized Veo 3 variant for rapid video creation and iteration, ideal for short-form content.
Qwen-Image 2512
Latest Qwen text-to-image model with enhanced quality, detail, and prompt adherence.
FLUX.2 [klein] 9B
FLUX.2 [klein] 9B Edit
Qwen-Image Edit 2511
Latest Qwen image editing model. Supports prompt-guided transformations with enhanced quality.
LTX Video 2.3 Fast
Speed-optimized LTX video model with extended duration support up to 20 seconds and portrait mode.
LTX Video 2.3
Latest LTX video model with sharper details, cleaner audio, and portrait support for professional content creation.
FLUX.1 Kontext [max]
Premium model for max editing performance, superior typography, and visual narrative consistency.
Qwen-Image Edit 2509
Advanced image editing model. Enhanced performance for fine-grained manipulation.
Seedance Lite
Versatile and efficient video model. Optimized for speed and ideal for short clips and rapid prototypes.
Qwen Image 2
Qwen Image 2 standard text-to-image model with strong prompt adherence and diverse style support.
Z-Image Turbo
Ultra-fast 6B parameter image model from Tongyi-MAI, optimized for near real-time generation.
HiDream O1 Image Dev
HiDream's distilled O1 Image Dev model — create, edit, and personalize images up to 2K in one native model.
Kling v1.6 Pro
Powerful video model for high-fidelity, imaginative content with complex character motion.
xAI TTS V1
xAI's text-to-speech model with speech tags for expressive delivery in 20+ languages.
Hunyuan Video 1.5
High quality open-source video model from Tencent Hunyuan.
FLUX.2 [klein] 4B
FLUX.2 [klein] 4B Edit
Veo 2
Legacy video model generating high-definition, long-form cinematic video content.
FLUX.1 Kontext [pro]
Pro-grade multimodal model for fast, iterative editing, style transfer, and consistency.
Wan 2.2
Advanced video model. Enhanced visual consistency and detail for high-resolution 1080p content.
FLUX 1.1 [pro] Ultra
Delivers ultra-high-res (up to 4MP) images with superior photorealism, detail, and speed.
ElevenLabs Multilingual V2
Multilingual text-to-speech with natural voice selection and stability controls.
FLUX 1.1 [pro]
Next-gen FLUX model. 6x faster with enhanced prompt adherence and top-tier quality for production.
Bria FIBO Edit Restore 1.0
Bria's one-shot photo restoration, trained on licensed data.
FLUX.1 [pro]
Flagship commercial T2I model offering superior prompt adherence, quality, and stylistic outputs.
FLUX.1 SRPO [dev]
12B flow transformer fine-tuned with SRPO for exceptional photorealism and polished composition.
Runway Aleph 2
Runway's next-generation video-to-video model for restyling and transforming existing footage from a prompt.
Runway Gen-4 Aleph
Runway's video-to-video model for restyling and transforming existing footage from a prompt.
Runway Gen-4 Turbo
Runway's speed-optimized image-to-video model for rapid, high-quality clip generation.
Qwen-Image
Open-source T2I model. Excels at high-res images from complex text, notable for clear, stylized text.
Recraft V3
Versatile image model for graphic design. Generates legible, stylized text and scalable vector art (SVG).
Stable Diffusion 3
Latest open-weight MM-DiT model. Major improvements in quality, prompt following, and text rendering.
FLUX.1 Krea [dev]
Open-weight model co-developed with Krea AI. Excels in photorealism and aesthetics.
FLUX.1 [dev]
Powerful, open-weight 12B image model. Excels in image quality, prompt adherence, and commercial use.
Minimax Video 01
Accessible video model. Generates engaging 720p clips from text with good visual consistency.
FLUX.1 Kontext [dev]
Open-weight, multimodal model for context-aware image editing. Excels at iterative edits via text.
Wan 2.1
State-of-the-art video model. Generates detailed, stylistically diverse clips with fluid motion.
FLUX.1 [schnell]
Ultra-fast, open-source model. Generates high-quality images in 1-4 steps for rapid prototyping.
Hunyuan Video
Impressive open-source video model. Excels at stable, coherent sequences with high visual quality.
Ray 2 Flash
High-speed variant of Ray 2, optimized for rapid video creation. Perfect blend of speed and quality.
Ray 2
Large-scale, state-of-the-art video model for stunning realistic and coherent 1080p motion.
BiRefNet v2
High-quality image background removal.
SeedVR2 Image Upscaler
ByteDance's SeedVR2 model for high-quality image upscaling.
ElevenLabs Sound Effects
Generate custom sound effects from text descriptions with duration and prompt control.
Grok Imagine Image 2.0
xAI's Grok Imagine Image 2.0 model for high-quality text-to-image generation and editing.
Seedream 5.0 Pro Layerize
Split a finished image into independently editable layers with Seedream 5.0 Pro Layerize.
Qwen-Image Edit 2511 Multiple Angles
Qwen Image Edit 2511 with the Multiple Angles LoRA. Re-renders the input image from a chosen camera angle.
Tripo v3.0
Latest 3D model for production-quality assets with superior textures and clean geometry.
Ideogram Remove Background
Fast, clean background removal from Ideogram.
Meshy V6
Meshy V6. High-fidelity 3D asset creation focusing on PBR, quad mesh, and face rigging.
Meshy V7
Meshy V7. Text, image, and multi-image to 3D with PBR, quad mesh, and optional rigging.
SeedVR2 Video Upscaler
Powerful video upscaling model from ByteDance, optimized for high-quality resolution and fidelity improvement.
Bria Video Background Removal
Remove video backgrounds for advanced video editing.
Clarity Creative Upscaler
Upscale images with high fidelity or creativity.
Hunyuan 3D v3.1 Pro
Latest Hunyuan 3D model with enhanced quality, multi-view input, and PBR material support.
Topaz Enhance
Enhance image quality with advanced upscaling.
Pixelcut Background Removal
High-quality image background removal from Pixelcut.
Topaz Video Upscaler
Professional-grade video upscaling solution utilizing Topaz AI technology for high-quality enhancement.
Recraft Vectorize Image
Convert raster images to vector graphics using Recraft.
Qwen-Image Layered
Split an image into layers using Qwen-Image Layered.
Topaz Wonder 3.5
Topaz's generative image upscaler, rebuilding detail with fewer repeated patterns.
Topaz Standard V2
Topaz Gigapixel precision upscaling, faithful to the original image.
Topaz Generative Enhance
Topaz Generative Enhance for enhanced details.
Sonilo Sound Effects 1.1
High-quality, commercial-use-safe sound effects from text with exact duration control.
Tripo v3.1
High-detail Tripo model for production assets with dense geometry and refined textures.
Gemini Omni Flash 1.1
Google's multimodal video model — 720p, 1080p, or 4K clips with synchronized native audio from text or a still image.
Topaz Starlight Precise 2.6
Topaz's generative video upscaler, rebuilding detail that the source never had.
FLUX.3
Frontier video model from Black Forest Labs with native audio, lipsync, and first/last frame control.
Meshy V5 Remesh
Dedicated Meshy 3D tool for remeshing models to optimize topology and reduce polygon count.
Uthana Animate
Retargets a Uthana motion — trained, or generated from a prompt or video — onto your character.
Recraft Crisp Upscale
Boost resolution while refining small details and faces.
Muse Image 1.0
Meta's Muse Image model with faithful instruction-following and precise multi-reference editing.
Bria Video Background Removal v3
Bria's v3 video background removal for advanced video editing.
Pixelcut Video Background Removal
Remove video backgrounds with Pixelcut for clean cutouts.
Rodin Gen-1.5
Advanced 3D generation model creating high-quality, textured T-pose avatars from a single image.
ESRGAN Upscaler
Very fast upscaling with good quality.
Topaz Transparency Upscale
Topaz upscaling that preserves the alpha channel end to end.
Image to SVG
Convert raster images to clean, scalable SVG vector graphics with fine-grained control over detail.
Lyria 3 Pro
Premium tier of Google's Lyria 3 with highest fidelity and compositional complexity.
OmniHuman 1.5
ByteDance's latest OmniHuman, bringing a still photo to life from audio with improved motion and expression.
Tripo Remesh
Rebuilds a high-poly 3D mesh as a low-poly model at a target face count, baking the original textures onto the result.
Sonilo Music 1.1
Licensed, commercial-use-safe music from a text prompt with exact duration control.
ACE-Step
Fast open-source music generation with lyrics alignment, remixing, and audio editing.
Meshy Rigging
Auto-rigs humanoid 3D models and optionally applies a preset animation, returning rigged GLB/FBX.
Wan 3.0
Alibaba's Wan 3.0 video model with up to 30s clips, 480p/720p/1080p output, and native audio.
Seed Audio 1.0
Bytedance Seed Audio TTS with preset voices and zero-shot voice cloning.
Inworld TTS 1.5 Max
Inworld's premium TTS model optimized for expressive game character voices.
Sora 2 Pro
Premium video version offering higher resolution (up to 1024p) and enhanced controls for pro projects.
FLUX.3 Video Upscale
FLUX.3-powered source-faithful video upscaler, up to 4K.
SAM 3.1 Image
Meta's SAM 3.1 for fast multi-object image segmentation.
FLUX.3 Extend
Extend an existing video clip with FLUX.3, continuing motion and native audio from the source ending.
Hunyuan 3D 3.0
Professional-grade 3D model optimized for high-quality, detailed assets with advanced features.
ESRGAN Video Upscaler
ESRGAN model for video upscaling, enhancing resolution and detail.
Tripo Retexture
Retextures an existing mesh from a text prompt or reference image, keeping geometry intact.
Tripo Rig
Auto-rigs a 3D mesh, returning a rigged GLB. Humanoid meshes come back with a Mixamo-named skeleton for direct use in game engines; other body plans use Tripo's own naming, which Mixamo cannot describe.
Tripo Segment
Segments a 3D mesh into parts, returning a part-grouped model for further editing.
Hunyuan Video Foley
Specialized model to automatically create and sync sound effects (foley) for video content.
Sonilo Video to Sound Effects 1.1
Synchronized, royalty-free sound effects timed to actions visible in a video.
Rodin Gen-2.5
Hyper3D Rodin Gen-2.5 pushes structural detail and surface quality further, with selectable quality.
Video Subtitles
Tripo P1
Game-ready Tripo model: a single image to clean low-poly 3D with PBR textures.
Sora 2
Next-gen video model generating long, high-fidelity 720p video with unparalleled narrative understanding.
Tripo Animate
Rigs a 3D mesh and applies a preset animation from the animation library, returning an animated GLB.
Meshy V5 Retexture
Specialized Meshy 3D tool for quickly retexturing imported meshes using text prompts.
VEED Subtitles
Hi3D Parts
Splits a 3D model into parts for editing, printing, or modular asset workflows.
Recraft Creative Upscale
Upscale for a sharper, cleaner, higher-resolution result.
Stable Audio
Open-source text-to-audio model for sound effects, field recordings, and instrument samples.
Lyria 3
Google DeepMind's next-generation music model with enhanced composition and vocal quality.
Trellis 2
Open-source, high-quality 3D model from Microsoft, leveraging a novel field-free sparse voxel structure.
Bria Video Increase Resolution
Bria's advanced AI technology designed to increase the resolution of video content efficiently.
Magi Distilled
Fast, efficient open-source I2V model. Animates still images using an autoregressive approach.
Hunyuan 3D v2 Mini
Lightweight, efficient 3D version optimized for less powerful hardware. Delivers good quality assets.
Recraft V4 Pro Vector
Recraft's pro V4 text-to-vector tier, with higher fidelity than the standard vector tier.
Recraft V4 Vector
Recraft's V4 text-to-vector model. Generates scalable vector art (SVG) directly from a prompt.
Gemini Omni Flash 1.1 Reference
Reference-guided video generation with up to 10 images and 3 short clips for consistent characters and style.
Google Virtual Try-On 1.0
Google Virtual Try-On: dress a person photo with a clothing product image.
FLUX.3 Video Upscale Creative
FLUX.3-powered video upscaler that adds detail, up to 4K.
Hunyuan 3D UV Unwrap
Rebuilds an existing mesh's UV layout so it is ready to texture. Geometry is kept as-is and the result is untextured.
Hunyuan 3D v3.1 Part
Splits a 3D FBX model into separate parts with Hunyuan 3D v3.1.
Topaz SDR to HDR
Topaz SDR-to-HDR conversion for grading and mastering.
Topaz Themis 2
Topaz motion deblur, faithful to the source footage.
Topaz Nyx HF
Topaz's precision video denoiser for pipeline-ready output.
Topaz Nyx XL
Topaz's video denoiser tuned for extreme noise.
Topaz Nyx Fast
Topaz's half-price video denoiser for high-volume work.
Topaz Nyx
Topaz's high-quality video denoiser, preserving texture.
Topaz Dust-Scratch V2
Topaz film clean-up for dust and scratches.
Topaz Recover 3 Restore
Topaz's generative restoration, rebuilding natural detail.
Topaz Denoise Max
Topaz's generative denoiser, rebuilding detail as it cleans.
Topaz Denoise Extreme
Topaz's most aggressive conventional denoiser.
Topaz Denoise Strong
Topaz denoising tuned for heavier noise.
Topaz Denoise Normal
Topaz general-purpose denoising at the source resolution.
Topaz Aion
Topaz's frame interpolator for extreme and complex motion.
Topaz Chronos
Topaz's frame interpolator tuned for real-world motion.
Topaz Apollo
Topaz's general-purpose frame interpolator, retiming footage up to 120fps.
Topaz Proteus Natural
Topaz's softer precision upscaler, tuned to look untouched.
Topaz Starlight Mini
Topaz's generative video upscaler tuned for archival restoration.
Topaz Rhea
Topaz's maximum-detail precision upscaler.
Topaz Theia Fine Tune Detail
Topaz's manually tuned precision upscaler, biased toward detail.
Topaz Iris
Topaz's precision upscaler specialised in recovering facial detail.
Topaz Proteus
Topaz's precision video upscaler, enhancing detail without inventing it.
Topaz Recover 3
Topaz's restorative upscaler, rebuilding natural detail in damaged images.
Topaz Artemis High Quality
Topaz's precision upscaler that denoises and sharpens as it enlarges.
Topaz Gaia 2
Topaz's animation-tuned precision upscaler, and its cheapest tier.
Topaz Dione DV
Topaz's deinterlacing precision upscaler for legacy tape sources.
Topaz Recovery V2
Topaz's upscaler for extremely low-resolution sources.
Topaz Astra 2
Topaz's creative video upscaler, reinventing detail for maximum visual impact.
Topaz Gaia CG
Topaz's precision upscaler for rendered and CG video.
Topaz Gaia HQ
Topaz's precision upscaler for refining already-clean footage.
Topaz Starlight Fast 2
Topaz's fastest and cheapest generative video upscaler.
Topaz Bloom Realism
Topaz's creative upscaler with its invented detail biased toward photorealism.
Topaz Standard MAX
Topaz's precision-leaning generative upscaler.
Topaz High Fidelity V3
Topaz's detail-preserving precision upscaler for professional photography.
Topaz Starlight HQ
Topaz's highest-quality generative video upscaler.
Topaz Starlight Sharp
Topaz's faster generative video upscaler with sharper output.
Topaz Bloom 2
Topaz's creative image upscaler, reinventing detail with an adjustable creativity dial.
Topaz Low Resolution V2
Topaz's precision upscaler tuned for small, compressed sources.
Topaz Text Refine
Topaz's precision upscaler that keeps text and hard shapes crisp.
Topaz Redefine
Topaz's prompt-guided generative upscaler.
Topaz CGI
Topaz's precision upscaler for rendered art and CG rather than photographs.
Hi3D v3.0
Hi3D v3.0 image-to-3D at 2048³ with PBR materials and face counts up to 5M.
VEED Lipsync 2.0
Production-quality lipsync that re-articulates a source video to match any new audio track.
LatentSync
Fast, affordable video-to-video lipsync that syncs mouth motion to any audio track.
sync.so Lipsync 2 Pro
High-quality realistic lipsync that preserves unique facial details from any new audio track.
sync.so Lipsync 3
sync.so's most powerful lipsync model, re-articulating a source video to any new audio track.
sync.so Avatar (Image to Video)
Turns a single still image into a talking character lip-synced to a voice track.
Uthana Auto-Rig Character
Uthana's automatic character rigging model, preparing an uploaded mesh with a skeleton.
Auto Caption
Wan 3.0 Reference
Reference-guided Wan 3.0 video with up to 10 images, 5 videos, and 5 audio files.
MiniMax Music 3
High-performance MiniMax music model for complete songs up to five minutes with structure tags and seed control.
FLUX Pro VTO
Virtual try-on from Black Forest Labs: dress a person photo with a garment reference.
SAM 3.1 Video
Meta's SAM 3.1 for multi-object video segmentation and tracking.
Sonilo Video to Music 1.1
Frame-synced, licensed music scored from video pacing, mood, and timing.
Sonilo Video Music 1.1
Mux frame-synced, licensed music onto any video; optionally keep original speech.
Sonilo Video Sound Effects 1.1
Add synchronized, royalty-free sound effects mixed into the finished video.
Rodin Bang
Segments a 3D mesh into parts with Hyper3D Bang!
Rodin Gen-2
Hyper3D Rodin Gen-2 delivers sharper geometry and cleaner textures from a single image or prompt.
Hi3D
Hi3D image-to-3D with multi-view support, PBR materials, and configurable face counts.
Hi3D Texture
Textures an existing Hi3D-compatible geometry mesh from a reference image.
FLUX.3 Keyframes
Guide FLUX.3 video generation with up to 10 keyframe images pinned across the clip.
Reve 2.0 Remix
Reve's remix model - blends one to eight reference images into a single new image, guided by a prompt.
Recraft V4.1 Vector
Recraft's V4.1 text-to-vector model. Generates scalable vector art (SVG) directly from a prompt.
Reve 2.0
Reve's flagship image model with best-in-class prompt adherence and text rendering, plus native editing.
Recraft V4.1 Pro Vector
Recraft's V4.1 Pro text-to-vector model. Generates scalable vector art (SVG) directly from a prompt.
Hunyuan 3D v3.1 Fast
Fast Hunyuan 3D model for rapid 3D generation with PBR material support.
Minimax Music V2.6
Latest MiniMax music model with native audio-visual generation capabilities.
Minimax Music V2.5
Updated MiniMax music model with improved vocal quality and arrangement control.
Stable Audio 2.5
Enterprise-grade audio generation with multi-part compositions and audio inpainting.
CassetteAI Sound Effects
Real-time sound effects generation in under 1 second, up to 30 seconds long.
SAM Audio Separate
Foundation model for audio source separation using text, visual, or temporal prompts.
ElevenLabs Text to Dialogue
Generate multi-speaker dialogue audio from structured text with per-speaker voice control.
ElevenLabs Voice Changer
Transform the voice in an audio recording to a different target voice.
CassetteAI Music
Ultra-fast music generation producing 3-minute tracks in under 10 seconds at 44.1 kHz.
Lyria 2
Google DeepMind's high-fidelity 48 kHz music model with fine-grained creative control.
Minimax Music V2
AI music generator producing complete songs with vocals, lyrics, and full instrumentation.
DeepFilterNet3
Real-time speech enhancement and noise suppression at 48 kHz full-band audio.
Kling Video to Audio
Generate synchronized sound effects, dialogue, and ambient audio from video content.
ElevenLabs Audio Isolation
Isolate clean voice from noisy recordings by removing background noise and music.
Mirelo SFX 1.6
Sound effects generation and editing with text-to-audio and audio inpainting.
ElevenLabs Music
Premium AI music generation from ElevenLabs with composition planning and section control.
Gemini TTS
Google's text-to-speech model with multi-speaker support and natural expressiveness.
Bytedance Seed 3D
Powerful 3D model focusing on high-quality objects from a single image. Adept at geometry & texture.
SAM 3D Align 1.0
Meta SAM 3D Align places body and object meshes into a spatially coherent scene from one image.
SAM 3D Body 1.0
Meta SAM 3D Body reconstructs accurate human body shape and pose from a single image.
SAM 3D Objects 1.0
Meta SAM 3D Objects reconstructs textured 3D geometry from a single real-world image.
OmniHuman
Advanced video model bringing a still image of a person to life using audio, producing expressive videos.
AI Avatar
Specialized model for creating realistic, audio-driven talking avatars with accurate lip-sync and expressions.
Tripo Turbo v1.0
Speed-optimized 3D generation model designed for rapid prototyping and fast generation times.
Framepack
Highly efficient, open-source I2V model that generates video by predicting the next frame.
Tripo v2.5
Incremental 3D update. Refined performance, improved mesh topology, and texture fidelity.
Hunyuan 3D v2
Powerful, open-source 3D model producing high-res, textured 3D objects from text or image inputs.
Minimax Video 01 Live
Specialized video model. Optimized for a dynamic, live-action feel with naturalistic camera work.
Trellis
Open-source 3D model creating high-quality objects with realistic materials and geometry from text.
MiniMax H3 Max Reference
Reference-guided MiniMax H3 Max video with multimodal image, video, and audio inputs at 480p or 768p.
No models match your search.