The newsletter of kiselbsthilfegruppe.de is now called rundbrief.ai

Subscribe
Menu

AI Tools · Category

AI Photo & Video Generation

The biggest category in the library: image and video models from Midjourney and Flux to Kling and Veo, audio with ElevenLabs and Suno, plus infrastructure like Replicate and fal.ai for building your own pipelines.

Tools & Workflows

32 tools
AI Photo & Video Generation primary tool

Midjourney

Develop image and video ideas with a strong visual signature.

Read our take

Midjourney is a core creative tool when a look needs to become high quality and visually distinctive quickly. It works well for moodboards, campaign visuals, style exploration and concepting before a team moves into production or finishing. For repeatable brand work it still needs clear references, consistent prompts and deliberate selection rather than isolated pretty images.

Visit website
AI Photo & Video Generation primary tool

Runway

Video and image generation for creative production flows.

Read our take

Runway belongs in workflows where images, text or existing footage need to become moving variants quickly. The platform is especially useful for ideation, edit experiments, B-roll, style tests and visual transitions. Its value appears when generation, editing taste and post-production are considered together.

Visit website
AI Photo & Video Generation primary tool

Krea

Creative suite for images, video and fast visual variants.

Read our take

Krea is practical when a team wants to test many visual directions side by side. It brings image, video and editing models into one interface, making it useful for exploration and creative direction. For final assets, the strongest outputs still need to be curated and finished carefully.

Visit website
AI Photo & Video Generation primary tool

Kling

Image-to-video and text-to-video for short cinematic clips.

Read our take

Kling is relevant when reference images or short prompts need to become dynamic clips. It works for product shots, social assets, motion studies and fast video ideas with relatively little setup. As with all video models, motion logic, hands, faces and edits need deliberate review.

Visit website
AI Photo & Video Generation usage-based

Replicate.com

Inference provider for image, video, audio and other AI models.

Read our take

Replicate.com is a useful place to try models without setting up infrastructure and later call them through APIs. For creative AI workflows this is helpful because many image, video, audio and upscaling models become easy to compare. In production, costs, runtimes, model versions and rights should be documented carefully.

Visit website
AI Photo & Video Generation usage-based

fal.ai

Fast inference APIs for generative media models.

Read our take

fal.ai is relevant when image, video, audio or 3D models need to be built into products with low latency and a clear API. The provider is especially useful for teams that want to compare models quickly, automate workflows and test generation beyond a web UI. Cost control, model choice and monitoring per workflow remain important.

Visit website
AI Photo & Video Generation usage-based

Runware

Inference platform for image generation, editing and media pipelines.

Read our take

Runware fits workflows where generative media is not only tested but connected to a product or internal pipeline through an API. The provider is interesting for image generation, editing, upscaling and fast variants through programmatic access. For teams, the key question is how reliably latency, cost and output quality fit their own use case.

Visit website
AI Photo & Video Generation primary tool

ElevenLabs

AI voices, voiceover, dubbing and audio for moving images.

Read our take

ElevenLabs is the natural building block when video needs voice, sound and dubbing in addition to visuals. It works for voiceover, localization, character voices, prototypes and fast audio versions. With voices, consent, disclosure and rights matter more than technical quality alone.

Visit website
AI Photo & Video Generation primary tool

Suno

Generate music ideas and songs from short prompts.

Read our take

Suno is helpful when videos, reels or prototypes quickly need a musical direction. Text ideas become songs, sketches and moods that can support edit rhythm and campaign concepts. For public use, licensing, plan terms and brand fit should always be checked deliberately.

Visit website
AI Photo & Video Generation primary tool

Figma Weave

Node-based AI generation for image, video and motion on one canvas.

Read our take

Figma Weave brings generative AI and professional editing into a node-based canvas where image, video, animation and VFX come together. It is strong when a look should not just be generated but branched, composited and iterated under control without switching tools. Since Figma Weave comes from the Weavy acquisition and currently runs separately from Figma, it is worth checking plan, rights and the integration announced for 2026.

Visit website
AI Photo & Video Generation primary tool

Higgsfield

Multi-model studio for image and video generation with character consistency.

Read our take

Higgsfield bundles many frontier image and video models into one studio, complemented by its own models such as Soul and Nano Banana. It is especially useful when character and style consistency across several visuals matters, for campaigns, storyboards or lookbooks. As with all multi-model platforms, model choice, cost and usage rights should be checked deliberately per use case.

Visit website
AI Photo & Video Generation key model

Flux Kontext Pro

Context-aware image editing for existing visuals.

Read our take

Flux Kontext Pro is interesting when an existing image needs targeted changes without losing the overall look. It fits iterations on products, people, scenes, backgrounds and campaign visuals. The workflow becomes strong when reference, desired change and quality control stay clearly separated.

Visit website
AI Photo & Video Generation key model

Flux Schnell

Fast text-to-image model for broad exploration.

Read our take

Flux Schnell is useful when many directions need to be screened quickly. It is less the final quality anchor and more an accelerator for early variants, style tests and prompt exploration. In teams, the model helps decide faster which visual idea deserves more work.

Visit website
AI Photo & Video Generation key model

Google Imagen 4

Google image model for high-quality text-to-image results.

Read our take

Google Imagen 4 belongs in the list because it is relevant for precise, high-quality image generation and visual tests. It is especially useful when text understanding, clear subjects and clean details matter. In creative workflows it should be compared with other models because every model has its own strengths and recurring failure modes.

Visit website
AI Photo & Video Generation key model

Seedream

ByteDance Seed models for images, references and video transitions.

Read our take

Seedream represents the ByteDance Seed family around image generation, editing and visual reference work. For video workflows, Seedance in the same ecosystem often becomes relevant because image-to-video and video references are stronger there. The building block is useful when existing frames, looks or character references need to become new visuals.

Visit website
AI Photo & Video Generation key model

Runway Aleph

In-context video editing for existing footage.

Read our take

Runway Aleph is relevant when existing videos need targeted changes rather than only new clips. It fits objects, lighting, style, scene variants and edit-specific corrections. For professional results, the key question is whether changes remain stable across multiple shots.

Visit website
AI Photo & Video Generation key model

Wan 2.2

Open-source video model for text and image-to-video.

Read our take

Wan 2.2 is an important open-source building block for teams that want more technical control over video generation. It is interesting for local setups, ComfyUI workflows, experiments and reproducible model tests. The advantage is openness; the effort is hardware, setup and workflow maintenance.

Visit website
AI Photo & Video Generation key model

Hailuo 2.0

MiniMax video generation for short realistic clips.

Read our take

Hailuo 2.0 is interesting when text or images should quickly become realistic-looking clips. It fits social visuals, storyboards, motion studies and fast variants for creative reviews. Realism does not automatically mean usability, so every output still needs checks for continuity and rights.

Visit website
AI Photo & Video Generation key model

Veo 3

Google video model for high-quality clips and audio context.

Read our take

Veo 3 is an important reference point for high-quality AI video. It is especially relevant when visual quality, motion and sometimes audio need to be considered together. For teams, Veo is also a benchmark against which their own toolchains and alternative models can be compared.

Visit website
AI Photo & Video Generation key model

Topaz Image Upscale

Image upscaling and enhancement for final assets.

Read our take

Topaz Image Upscale belongs near the end of a workflow rather than the beginning. It helps prepare good images for higher resolution, sharper detail or technical delivery. It is especially useful when generated visuals need to look clean in presentations, print, web headers or campaign formats.

Visit website
AI Photo & Video Generation key model

Topaz Video Upscale

Video upscaling and enhancement for better delivery quality.

Read our take

Topaz Video Upscale is a finishing tool for clips that already work creatively but need stronger technical quality. It can help with resolution, sharpness, denoising and stability. This matters for AI video because generation models often deliver good ideas but not always final delivery quality.

Visit website
AI Photo & Video Generation key model

Magnific Image Face Enhancer

Creative upscaling and face enhancement for image assets.

Read our take

Magnific is relevant when an image should not only become larger but look clearer, more detailed and more premium. Creative enhancement can make a strong difference for faces, textures and campaign visuals. At the same time, overprocessing needs attention because too much enhancement quickly looks artificial.

Visit website
AI Photo & Video Generation key model

All Black Forest Labs Models

BFL model family around Flux for image generation and editing.

Read our take

The Black Forest Labs models are an important reference point for modern image generation. In practice, Pro, Schnell, Dev and Kontext variants should be compared by cost, speed and quality separately. Teams get value when they do not blindly use one model but choose the variant that fits the workflow.

Visit website
AI Photo & Video Generation key model

Flux Fine Tunes

Adapted Flux models for repeatable brand and style worlds.

Read our take

Flux fine tunes become interesting when a style, product world or character should not need to be recreated from scratch in every prompt. They can create consistency and speed up work on recurring visuals. The effort is most worthwhile when enough strong training examples and clear quality criteria exist.

Visit website
AI Photo & Video Generation key model

Kontext Fine Tunes

Specialized Kontext workflows for repeatable image editing.

Read our take

Kontext fine tunes represent workflows where editing should be repeatable across similar visuals rather than one-off. This is interesting for product variants, series, campaign worlds or consistent people and object edits. The important part is keeping changes not only attractive but stable and traceable.

Visit website
AI Photo & Video Generation key model

Flux Dev Trainer

Training workflows for custom Flux Dev LoRAs and style models.

Read our take

A Flux Dev Trainer is relevant when custom LoRAs, product looks or style models need to be built. It helps turn curated examples into repeatable visual building blocks. Data quality, clean captions and a test set are decisive for knowing whether the model really generalizes.

Visit website
AI Photo & Video Generation key model

3D Models Collection

Collection for 3D generation, assets and experimental workflows.

Read our take

3D models belong in the list because image and video work increasingly overlaps with spatial assets, product objects and scene variants. A collection is practical for comparing models, demos and approaches before committing to a pipeline. Maturity varies widely, so early testing matters more than big promises.

Visit website
AI Photo & Video Generation deep dive

Hugging Face

Model hub for open and commercial image, video and audio models.

Read our take

Hugging Face is the deep-dive place when models should be understood and compared rather than only used. Teams find model cards, demos, weights, Spaces and community signals there. For production use, licensing, provenance, safety boundaries and technical requirements matter a lot.

Visit website
AI Photo & Video Generation deep dive

Stable Diffusion Web

Stability AI models and APIs for image generation and editing.

Read our take

Stable Diffusion Web represents entry into the Stability ecosystem through web tools, APIs and model variants. It is useful when teams want to connect image generation with more control, model understanding or their own technical integration. The deep dive is especially worthwhile when prompting, control, inpainting and model choice work together.

Visit website
AI Photo & Video Generation deep dive

Automatic1111 Web UI

Classic Stable Diffusion web interface for local experiments.

Read our take

Automatic1111 Web UI is a well-known entry point for local Stable Diffusion workflows. It fits experiments, extensions, inpainting, upscaling and classic prompt setups. For new teams it can be more approachable than some node workflows, but it is less flexible than highly modular setups.

Visit website
AI Photo & Video Generation deep dive

ComfyUI

Node-based interface for controlled image and video pipelines.

Read our take

ComfyUI is strong when generation is treated as a reproducible pipeline. Nodes, workflows and model combinations make complex image and video experiments more traceable. The entry is more technical, but for teams with repeatable creative pipelines it is often much more powerful.

Visit website
AI Photo & Video Generation deep dive

ThinkDiffusion

Cloud environment for ComfyUI, Automatic1111 and open GenAI workflows.

Read our take

ThinkDiffusion is interesting when open image and video workflows should be used without setting up local hardware. It fits teams that want to test, share and run ComfyUI or Automatic1111 more practically. Its value is making deep-dive workflows more accessible without giving all control to a simple consumer app.

Visit website

Rundbrief

Good tools. Straight to your inbox.

New tools and the AI developments that matter. Free, in our German-language newsletter.