Image and Video Generation AI News — 2026-09-12
Open-source contenders—especially China’s Z-Image, Qwen-Image, and MiniMax-H3—are closing in on closed-source models in blind evaluations. Meanwhile, Google is mounting a broad offensive with Nano Banana, Veo, and Gemini Omni, while OpenAI’s Sora shows signs of slowing down following model removals from its pricing page and the departure of its lead.
Image and Video Generation AI News Roundup — 2026-09-12
Hey! This time I did a serious investigation of the image and video generation AI space across six platforms: Reddit, X, YouTube, Bluesky, Lemmy, and Pinterest 🔥 Bottom line: on the closed-source side, Google is going full throttle on both image and video, while OpenAI’s Sora seems to be losing momentum 😳. On the open-source side, Chinese players (Qwen, Z-Image, MiniMax, Wan) are getting genuinely close in blind evaluations, making this an incredibly exciting development 🔥🔥 That said, the X research was a complete miss: it surfaced zero image/video-generation-AI posts, so I’m being upfront about that major gap 💦 Across every platform, though, concern over whether content is AI-generated—and whether it can be trusted—was hugely active. Surprisingly, it spread more widely than discussions of model performance 👀
Across platforms
- China has completely taken center stage in open source 🇨🇳: Across YouTube, Bluesky, and Lemmy independently, Qwen-Image, Z-Image (Turbo), MiniMax (H3/Hailuo 3.0), and Alibaba Wan are being praised as the “new champions.” On Bluesky’s blind leaderboard, Qwen-Image-3.0-Pro is reportedly just 92 Elo points behind GPT Image 2 (high), which is genuinely impressive (https://bsky.app/profile/oludai.bsky.social/post/3mv5ywkauvt2t).
- Google is attacking on every front across image and video: YouTube has reviews of Veo 3.1 (https://www.youtube.com/watch?v=sTOFXi2eY_k) and Teacher's Tech’s Nano Banana 2 introduction (https://www.youtube.com/watch?v=nikIxHM_AQY). Bluesky even reported a Windows desktop app with “Gemini Omni” integration (https://bsky.app/profile/itmatterss.bsky.social/post/3mvayanva5c2n). The all-out approach looks seriously strong 💪
- “Can you spot the AI?” content is genuinely popular: Identification threads on Reddit (r/isthisAI) (https://www.reddit.com/r/isthisAI/comments/1wbspxo/) and Pinterest infographics about deepfake/AI-voice detection are independently gaining traction. Authenticity anxiety has become a cross-category theme. It is quietly surprising that it travels farther than model-performance news.
- “Unlimited” and affiliate marketing claims are viewed skeptically everywhere: Reddit comments warn that sites advertising “unlimited” are often borderline scams (#4, r/generativeAI), while circuitai.bsky.social repeatedly posted identical WAN 3.0 spam copy more than 15 times a day on Bluesky. Users are appropriately skeptical of promotional claims 😤
- Mainstream users do not really distinguish open from closed models: In Reddit tool-advice threads, Qwen, Flux, Seedream, Veo, and GPT Image are discussed side by side as a list of usable models. Almost no comments explicitly distinguish between them. There is a notable gap between the brief’s two-axis framing and how users actually think.
Platform by platform
Reddit — 12 threads collected (short of the target of roughly 20). The most successful post was not about a new model but about using GPT-6 to control a mouse and create fake hand-drawn timelapses (r/antiai, 7,146 points, https://www.reddit.com/r/antiai/comments/1w81kdd/), reflecting anger at AI “deception.” Multiple threads independently mentioned uncensored, NSFW-capable unified platforms such as WildOwl.ai, cyberbara, and Tensor.art, suggesting demand.
X — A complete miss 😭 The 40 collected posts were all about Yemen, the $SONG memecoin, Chelsea FC, and similar topics; there were zero posts about image or video generation AI. The search terms (trending topics) were misaligned with the subject, so next time it needs a fresh search using specific names such as Midjourney, Sora, and Runway.
YouTube — The most information-rich platform this time ✨ On the closed side, GPT Image 2/2.5 reportedly leads the Artificial Analysis Image Arena (Elo 1,339, the largest-ever gap between first and second place, https://www.youtube.com/watch?v=DY6TkI8XtkQ); Veo 3.1 improves lip-sync; and ByteDance Seedream 5.0 Pro is said to overwhelm existing models (https://www.youtube.com/watch?v=WpcxwSNT3X4). On the open side, Black Forest Labs FLUX.2 is discussed as a potential “new champion” (https://www.youtube.com/watch?v=YQuTkPVkCS8), Alibaba Z-Image Turbo is called the KING of local models (https://www.youtube.com/watch?v=K0FgDU-9uik), and Lightricks LTX-2 is even described as the “holy grail of open-source video.”
Bluesky — Sora may be seriously losing momentum 👀 Alongside a post detecting the quiet removal of sora-2 and sora-2-pro from the pricing page (https://bsky.app/profile/ctxwindow.kynth.studio/post/3muxriifygl23), there was a scoop that former Sora lead Bill Peebles is launching an independent startup with major Hollywood figures (https://bsky.app/profile/theinformation.com/post/3mv6myz7pa52n). On the open side, Chinese players—Qwen, Z-Image, MiniMax, WAN 3.0, and SenseNova-U1.5—dominate discussion.
Lemmy — Smaller in scale and effectively dependent on curation by a single user, Even_Adder 😅 Still, the biggest closed-source story is genuinely startling: Midjourney reportedly pivoted from image generation to a “medical spa business” using ultrasound CT scans (https://www.theregister.com/ai-and-ml/2026/06/18/midjourney-pivots-from-ai-image-generation-to-body-scanning-medical-spa-where-patients-bathe-in-golden-light/5258429). It is being mocked as “Theranos all over again.” On the open side, the MiniMax-H3 ecosystem is exploding, with more than seven ComfyUI and quantization tools released in a week.
Pinterest — All 50 pins were checked visually 👀 Closed-source commercial tools—Runway, Pika, Kling AI, Luma, Canva, Hailuo, and Synthesia—appear in almost every ranking infographic, while almost no open-source pins were found. There is some real news, including official Google Veo 2/3 promotional materials and a ByteDance–Hollywood (MPA) copyright agreement (#31), but most content consists of mass-produced SEO/affiliate thumbnails. One pin (#44) claimed that more than 20% of YouTube is AI-generated.
What to watch
- Whether Sora’s slowdown is real — Bluesky’s detection of models removed from the pricing page (https://bsky.app/profile/ctxwindow.kynth.studio/post/3muxriifygl23)
- How much further Qwen-Image-3.0-Pro can close the gap with GPT Image 2 — Bluesky blind-evaluation leaderboard (https://bsky.app/profile/oludai.bsky.social/post/3mv5ywkauvt2t)
- The pace of MiniMax-H3 ecosystem growth — Lemmy, ComfyUI-related tools (https://github.com/sepiablue-ai/ComfyUI-MiniMax-H3-W4A4-VSA)
- Further developments in Midjourney’s medical-spa business — Lemmy, TheRegister article (https://www.theregister.com/ai-and-ml/2026/06/18/midjourney-pivots-from-ai-image-generation-to-body-scanning-medical-spa-where-patients-bathe-in-golden-light/5258429)
- The outcome of the “local champion” contest between FLUX.2 and Z-Image Turbo — YouTube (https://www.youtube.com/watch?v=YQuTkPVkCS8, https://www.youtube.com/watch?v=K0FgDU-9uik)
- Consolidation among uncensored, NSFW unified platforms such as WildOwl.ai — Reddit (https://www.reddit.com/r/AItips101/comments/1w9yftp/)
Recommendations
- For the next X investigation, search specific names such as Midjourney, Sora, Runway, Kling, Flux, and ComfyUI instead of trending terms.
- Continue following Bluesky accounts that monitor pricing pages to verify Sora’s trajectory.
- The clearest metric for tracking the open/closed performance gap is the Elo trend for Qwen-Image-3.0Pro versus GPT Image 2.
- Midjourney’s medical-business pivot is worth following as a business-strategy story rather than purely as technology news.
- Expand Reddit search terms next time so collection reaches the target of 20 threads rather than 12.
- Do not trust Pinterest ranking infographics or Bluesky affiliate-post floods in isolation; only use them when corroborated on other platforms.
Data quality
X contained zero image/video-generation-AI posts due to poor search-term selection, so it effectively provides no data for this subject. Reddit finished with 12 threads, below the target of roughly 20, and its content skewed toward tool recommendations rather than news. Pinterest source data had no open/closed categorization and required manual reclassification. Lemmy is a small platform effectively dependent on curation by one account, Even_Adder. Because YouTube pages are JavaScript-rendered, WebFetch could not retrieve page text, and view/subscriber counts could not be obtained even through oEmbed. Bluesky produced repeated 403 errors for some keywords (Runway, Kling, Black Forest Labs), so collection was abandoned for those; WAN 3.0 results were also saturated with spam-account reposts and need to be discounted.
Platform summaries
Reddit — Image and Video Generation AI News
Where
The 12 collected threads came from the following nine subreddits.
| Subreddit | Members | Collected threads |
|---|---|---|
| r/isthisAI | 408,823 | 1 |
| r/antiai | 322,973 | 1 |
| r/GetNoted | 298,488 | 1 |
| r/generativeAI | 152,415 | 3 |
| r/aiwars | 167,809 | 1 |
| r/hatethissmug | 134,660 | 1 |
| r/aivideomaking | 6,823 | 2 |
| r/GoogleFlow | 1,884 | 1 |
| r/AItips101 | 1,299 | 1 |
Large subreddits debating generative AI itself (r/antiai, r/aiwars, r/isthisAI, r/GetNoted) were prominent, while mid-sized tool-introduction and recommendation communities (r/generativeAI, r/aivideomaking) conveyed actual tool trends.
What people say
- 【#6, r/antiai, 7,146pt / 1,025 comments, 2026-09-05】A post claimed that an LLM (described as GPT-6) was given mouse and keyboard control, enabling it to “hand-draw” on a digital canvas and fake even a timelapse. The top comment (2,511pt, u/Nayutantantan) expressed anger: “I don’t even feel like recording timelapses of my own work anymore, because fake timelapses are becoming normal.” This single thread received by far the most engagement in the collection. https://www.reddit.com/r/antiai/comments/1w81kdd/
- 【#11, r/GetNoted, 4,470pt / 325 comments, 2026-09-10】An example of an X post using an AI-generated image—containing garbled text such as “Handloags” in a schedule graphic—became a source of ridicule. Top comments mocked the roughness of the generated image itself. https://www.reddit.com/r/GetNoted/comments/1wcw3n9/
- 【#1, r/AItips101, 145pt / 83 comments, 2026-09-07】“2026 Uncensored AI Image Generation Rankings.” WildOwl.ai ranked first, presented as offering pay-as-you-go access to more than 35 models including Qwen Image, Flux, Seedream, Wan, Kling, Seedance, Veo, and GPT Image, with no monthly fee and non-expiring credits. One commenter (u/reliablemicrophone84, 4pt) agreed, saying filters elsewhere had wasted their credits. Another (u/Pretend_Program_3696, 3pt) recommended Perchance/Uncensored.ai instead, so opinions were divided. https://www.reddit.com/r/AItips101/comments/1w9yftp/
- 【#8, r/aivideomaking, 4pt / 21 comments, 2026-09-09】In response to someone seeking free or credit-based photorealistic image/video tools, users listed OpenartAI, fal, dzine, pixverse, Higgsfield, and RunningHub (which supports both OSS and closed models via ComfyUI). u/zeroludesigner added that cyberbara’s Seedance 2.5 is uncensored and supports real faces. https://www.reddit.com/r/aivideomaking/comments/1wbxpdf/
- 【#3, r/generativeAI, 0pt / 7 comments, 2026-09-11】A thread seeking browser-only NSFW-capable image/video generation tools. u/MisterZan25 organized recommendations for RunDiffusion/Tensor.art (images and character consistency), PornGen/PornJourney, and Pika Labs mirror services (video). u/frighten (1pt) recommended the iPhone-only app “PixelForge” as on-device and uncensored. https://www.reddit.com/r/generativeAI/comments/1wdkbxv/
- 【#4, r/generativeAI, 1pt / 11 comments, 2026-09-07】In response to “What is the cheapest unlimited 1080p image-to-video service?”, Artlist.io (€550/year) was suggested. However, multiple comments criticized “unlimited” claims: u/ai_art_is_art (3pt) said, “Sites that advertise ‘unlimited’ are usually borderline scams—you put in $550, only get $200 worth of output, then wait 24 hours.” https://www.reddit.com/r/generativeAI/comments/1wa4k88/
- 【#5, r/aivideomaking, 16pt / 38 comments, 2026-09-07】Asked how to create a free 45-second video using only a phone/iPad, u/NewLifeWares (2pt) soberly replied that 45 seconds is effectively a data-center-scale demand; free users can realistically get only one or two five-second clips with watermarks and stitch them together. u/GuessingEngineer (2pt) joked that anyone who could do it in one shot would rule the world. The thread makes visible both demand for free, high-quality, long-form video and how unrealistic that combination remains. https://www.reddit.com/r/aivideomaking/comments/1w9h5a9/
- 【#7, r/GoogleFlow, 6pt / 2 comments, 2026-09-10】A bulk-generation Chrome extension for Google Flow (Veo), “AutoFlow,” was released. It supports batch generation, continuation from the last frame of a previous clip, and automatic editing of narrated videos with audio/subtitles, using a free tier plus Pro/Ultra subscriptions.
- 【#12, r/aiwars, 0pt / 14 comments, 2026-09-08】A discussion of likeness and advertising-use risks when an AI-generated person accidentally resembles someone real. u/Upstairs-Mammoth-509 (3pt) noted that accidental matches with unknown ordinary people are the hardest cases; companies can do little beyond reverse-image search and hope, highlighting lagging legal frameworks. https://www.reddit.com/r/aiwars/comments/1wafe86/
- 【#10, r/hatethissmug, 59pt / 20 comments, 2026-09-12】Different levels of intensity within anti-AI circles in response to a video that criticized generative AI while using it. u/AprilsStuff (22pt) wrote: “I’m anti-AI, but the environmental-impact argument is exaggerated, and saying ‘AI will naturally disappear’ is unrealistic.” Even anti-AI discourse has internal disagreement. https://www.reddit.com/r/hatethissmug/comments/1wdxv6k/
- 【#2, r/isthisAI, 0pt / 31 comments, 2026-09-09】A thread judging whether a video was AI-generated or live-action. A necklace disappearing and reappearing (u/smalltiddygothb, 16pt / u/EmergencyWaffleBox, 14pt) was considered decisive evidence. u/BouncingSphinx (5pt) noted that the video being exactly 15 seconds also felt like a current AI-video limit. Technical limitations such as length and inconsistencies are becoming standard cues for ordinary users. https://www.reddit.com/r/isthisAI/comments/1wbspxo/
Signals
- What is growing: References to uncensored, NSFW-capable unified platforms (WildOwl.ai, cyberbara, RunDiffusion, Tensor.art, PixelForge, and others) arise independently across multiple threads. Demand is strong for both access to the latest models and relaxed censorship.
- What is growing: The emergence of adjacent ecosystems that complement official tools, such as the third-party AutoFlow automation tool for Google Flow (#7).
- What users dislike: Subscriptions advertising “unlimited” access (#4, including Artlist.io) are consistently treated with suspicion.
- What users dislike: Even within anti-AI communities, arguments emphasizing environmental cost (#10) are criticized from within as exaggerated, showing anti-AI rhetoric is not monolithic.
- What was surprising: Of the 12 collected threads, the highest engagement came not from new-model releases but from an anti-AI exposé about faking hand-drawn timelapses (#6, 7,146pt) and a backlash-provoking AI image (#11, 4,470pt). On Reddit, AI deception and trustworthiness spread far more widely than product news.
- What was surprising: The open-source/closed-source distinction itself is not central to user interest. Most threads list models from both groups—Qwen, Flux, Seedream, Wan, Kling, Seedance, Veo, GPT Image, and others—side by side as available options, with almost no comments explicitly drawing the distinction.
Limits
- Only one search term, "Image and Video Generation AI News," was used, yielding 12 threads rather than the expected target of roughly 20.
- Collection relied on material pre-gathered by a worker using a headless browser, so additional Reddit searches and browsing could not be conducted directly in this run because Reddit blocked WebFetch and search-based access.
- Many collected posts were tool-recommendation Q&A threads rather than news; they reflect actual community tool use more than news trends. No specialist news threads clearly discussing closed-source and open-source developments were found.
- r/AItips101 (1,299 members) and r/GoogleFlow (1,884 members) are extremely small subreddits, so insights from them should be viewed as individual voices rather than representative of the broader community.
X
X — Image and Video Generation AI News Trends
Accounts
The 40 collected posts came from 40 separate accounts, one post per account, with no recurring accounts. None specializes in image/video-generation AI such as Midjourney, Stable Diffusion, Sora, Runway, Kling, Veo, ComfyUI, or Flux. Accounts clustered by search term.
- Yemen/Saudi (Middle East and OSINT): @tleilax___, @Ahmed175143, @IranTimes72, @citrinowicz, @Osinttechnical, @c14english, @Rusia_HD, @MerruX — OSINT/news accounts covering attacks on Saudi oil infrastructure and the Yemen conflict.
- $SONG (Cardano memecoin): @LagyADA, @EAjdarovic, @JureKaramarko, @SintoAndrej — small speculative/promotional accounts.
- Charlie Kirk (first-anniversary posts): @tivent_1, @TruthNetwork24, @mscaradapawjob, @rokkstarrrrrrr — political commentary and meme reactions.
- Britain (UK politics): @brigantia__, @Steven41849941, @SirJBritain, @BROKENBRITAIN0 — nationalist posts about immigration and public safety.
- Anthropic (AI economics and security): @restitutorII, @BrianRoemmele, @walterkirn, @claudecode84 — Anthropic economic-scenario forecasts, Claude Code open-sourcing, and AI criticism. These concern LLM agents and AI society rather than image/video generation.
- Apple/Samsung (smartphone arguments): @FlossyCarter, @cagriozturkish, @appleinteligen, @BrandonKHill — fan disputes around foldable phones.
- Sony (games industry): @JayDubcity16, @Cris_Uther, @treeblazah, @ScreenCraveHQ — gossip about Hideo Kojima and Spider-Man films.
- "one ai os" (AI startup self-promotion): @de1lymoon, @neda_sefati, @harvey_control, @Iamanusbutt — networking posts from AI-agent/business-OS startups, without image/video-generation discussion.
- Chelsea (football): @declanstar1, @TheHateCentra, @Lea_EFC, @chelcatastr0phe — Chelsea FC match content.
There were no repeated posters sufficient to distinguish one major viral account from several smaller recurring contributors. @rokkstarrrrrrr (9,404 likes), @TheHateCentra (19,654 likes), and @Cris_Uther (17,095 likes) received the highest one-off engagement, but all were off-topic.
Posts
There were zero posts directly mentioning image/video generation AI. The closest—but still off-topic—posts are listed for reference.
-
#21 @restitutorII — 458 likes · 69 reposts · 41 replies · ~57,000 views · 2026-09-10
https://x.com/restitutorII/status/2097918742571139093"Anthropic in its report imagines 3 possible futures for the American economy in 2030 based on the power and adoption of AI..."
→ Anthropic economic-scenario predictions. This concerns generative AI broadly, not image/video-generation trends. -
#24 @claudecode84 — 2,585 likes · 268 reposts · 24 replies · ~741,000 views · 2026-09-10
https://x.com/claudecode84/status/2097959830942367747“Anthropic’s CEO himself is open-sourcing his entire ‘Claude Code environment.’”
→ Claude Code open-sourcing. This is about coding AI, not image/video-generation models. -
#33 @de1lymoon — 49 likes · 6 reposts · 9 replies · ~3,500 views · 2026-09-11
https://x.com/de1lymoon/status/2098428267158032841"Grok Bot + Kimi K3 is where an AI agent starts turning into an operating system..."
→ AI-agent orchestration, unrelated to image/video generation. -
#22 @BrianRoemmele — 792 likes · 187 reposts · 105 replies · ~100,000 views · 2026-09-10
https://x.com/BrianRoemmele/status/2098024115231924568
→ A meme-style post criticizing Anthropic and discussing AI threats, without image/video-generation content. -
#23 @walterkirn — 723 likes · 91 reposts · 39 replies · ~28,000 views · 2026-09-10
https://x.com/walterkirn/status/2098059171325427949
→ Likewise, in the context of criticizing Anthropic/media, without mentioning image or video generation.
These five are the full set surfaced by the keyword “AI,” and none can be treated as news about either closed-source or open-source image/video generation AI.
Signals
- X Explore’s top ten trends in this session—likely Croatia-based given the “Trending in Croatia” label—were “Yemen,” “Saudi,” “$SONG,” “Charlie Kirk,” “Britain,” “Anthropic,” “Apple,” “Sony,” “one ai os,” and “Chelsea.” None involved image/video-generation terms such as Midjourney, Sora, Stable Diffusion, Runway, Kling, or Veo.
- The presence of “Anthropic” and “one ai os” suggests active discussion on X around AI agents, LLMs, and their economic impact, but that is a separate topic cluster from image/video generation AI.
- None of the 40 collected posts mentioned image/video AI tools, releases, demo videos, or comparisons.
Limits
- The search terms—"Yemen", "Saudi", "$SONG", "Charlie Kirk", "Britain", "Anthropic", "Apple", "Sony", "one ai os", and "Chelsea"—were X Explore trends for that day, not terms targeting image/video generation AI such as Midjourney, Stable Diffusion, Sora, Runway, Kling, Veo, ComfyUI, or Flux. Consequently, none of the 40 posts contained relevant news.
- As a result, this stage could not produce the brief’s requested insights—ten open-source and ten closed-source items—from X data. For this topic, no meaningful collection was effectively performed; the issue is not that X had nothing to say, but that topic-aligned searches were not conducted.
- Future runs must recollect data using names of image/video AI tools and companies: Midjourney, Stability AI, Runway, Pika, Luma, Kling, Sora, Google Veo, Flux, ComfyUI, and similar terms.
YouTube
YouTube — Image and Video Generation AI News (Closed/Open Source)
Channels
- There's An AI For That — An AI-tool introduction channel that posted a FLUX.2 review.
- AI Search — A new-model news channel covering Sora 2 and HunyuanVideo 1.5.
- Theoretically Media — A specialist channel with a strong focus on AI video generation and a detailed Veo 3.1 review.
- Teacher's Tech — A practical-tools channel with a Nano Banana 2 review.
- Codebreakers — An open-source AI model review/explainer channel featuring MiniMax H3.
- Bijan Bowen — A serious local-AI image-generation channel testing Z-Image Turbo.
- Benji's AI Playground — A ComfyUI/local-generation tutorial channel explaining Ideogram 4.0.
- Curious Refuge — An established AI filmmaking education channel reviewing Midjourney V8.
- RandomAI — A new-model-news and comparison channel explaining GPT Image 2.5.
- Standarity / The AI Filmmaking Advantage — News and comparison-testing channels covering the Seedream 5.0 Pro announcement and full test.
- MattVidPro AI — Mentioned through first-impressions videos on the Runway Gen-4 line.
Videos
- “Is FLUX 2.0 The New Open Source King?” — There's An AI For That — https://www.youtube.com/watch?v=YQuTkPVkCS8 — Tests whether Black Forest Labs’ FLUX.2 could become the new open-source image-generation champion.
- “We have a new #1 open-source AI video generator!” (2025-11-25) — AI Search — https://www.youtube.com/watch?v=6EQP8-D37bs — A HunyuanVideo 1.5 review and ComfyUI installation guide, presented as workable even on low-VRAM systems.
- “Sora 2… wtf” — AI Search — https://www.youtube.com/watch?v=El-G4cO4x4I — Compares OpenAI Sora 2 with Veo 3, Kling 2.5, and Hailuo 02, noting improvements in audio synchronization and physical behavior.
- “Veo 3.1 Is More Powerful Than You Realize!” (2025-10-16) — Theoretically Media — https://www.youtube.com/watch?v=sTOFXi2eY_k — Explains Google Veo 3.1 improvements in lip-sync, character consistency, and 1080p upscaling.
- “The New Gemini Image Generator is Insane (Nano Banana 2)” — Teacher's Tech — https://www.youtube.com/watch?v=nikIxHM_AQY — Reports that the Gemini app’s default image generator was replaced with Nano Banana 2.
- “MiniMax H3: The New Open-Source Video Champ?” (2026-08-03) — Codebreakers — https://www.youtube.com/watch?v=3R5GROj8Tcg — Introduces the open-weight, omnimodal video model MiniMax H3 (Hailuo 3.0), which can generate videos with native stereo audio in one pass.
- “Alibaba Z-Image Turbo Test – The New KING of LOCAL Image Models” — Bijan Bowen — https://www.youtube.com/watch?v=K0FgDU-9uik — Tests Apache-2.0-licensed Z-Image Turbo locally and calls it the KING.
- “Ideogram 4.0 Released! The New Open Model That FINALLY Gets Text Right” (2026-06-04) — Benji's AI Playground — https://www.youtube.com/watch?v=HY_e5CZfJfQ — Reports that Ideogram’s first open-weight model topped Design Arena and finally achieved accurate text rendering.
- “I Tried Midjourney 8… Here's the Truth” (2026-03-19) — Curious Refuge — https://www.youtube.com/watch?v=i9qV9-2tzEA — A candid capability review of Midjourney V8.
- “GPT Image 2.5 Is AMAZING: Full Breakdown of OpenAI's New Image Model” — RandomAI — https://www.youtube.com/watch?v=DY6TkI8XtkQ — Explains the GPT Image 2/2.5 family and improvements in text rendering and multilingual support.
- “ByteDance Seedream 5.0 Pro: The Announcement, Read and Highlighted” — Standarity — https://www.youtube.com/watch?v=g05iPef37Qs — Summarizes Seedream 5.0 Pro’s announcement, including layer-level editing.
- “🚨BREAKING: Seedream 5.0 Pro Just SMOKED Every Image Model (Fully Tested)” (2026-07-16) — The AI Filmmaking Advantage — https://www.youtube.com/watch?v=WpcxwSNT3X4 — Fully tests Seedream 5.0 Pro against other models and judges it to outperform the field.
Signals
Closed-source developments
- OpenAI GPT Image 2 / 2.5 reportedly leads Artificial Analysis Image Arena with an Elo score of 1,339 and the largest first-versus-second-place gap in the arena’s history as of September 2026.
- Google Veo 3.1 and Nano Banana Pro (Gemini 3 Pro Image) are strengthening both video and image offerings, with lip-sync, 1080p upscaling, and layer editing signaling a clear push toward professional use.
- ByteDance Seedance 2.0 (2026-02) and Seedream 5.0 Pro (around July 2026) are rapidly increasing their presence through multimodal inputs—text, image, video, and audio references.
- Several videos claim xAI Grok Imagine 1.5 exceeds Veo and Sora in benchmarks, making it a topic of conversation around June 2026.
- Existing players such as Runway Gen-4/4.5 and Midjourney V8 (V8.1, V8.2) continue to update, generating many “defending the throne” review videos.
Open-source developments
- Multiple reviews describe Black Forest Labs FLUX.2 as the “new champion” among high-quality local image models.
- Tencent HunyuanVideo 1.5 supports ComfyUI natively and is becoming a new standard for open-source video generation that runs in low-VRAM environments.
- Alibaba/Tongyi-MAI’s Z-Image and Z-Image Turbo were released under Apache-2.0 and are widely described as the “new king of local models.”
- Ideogram 4.0 launched as the company’s first open-weight model, reaching first place in Design Arena and drawing attention for accurate text rendering.
- MiniMax H3 (Hailuo 3.0) rapidly emerged in August 2026 as an open-weight omnimodal video model capable of running on personal GPUs such as the RTX 3090, with ComfyUI videos appearing in quick succession.
- Lightricks LTX-2 is described as achieving open-source 4K, 50fps, audio-synchronized video generation—the “holy grail of open-source video.”
- Alibaba Wan 2.5/2.6 continues to be covered as an open-source video family supporting text, image, audio, and video inputs.
Limits
- YouTube search pages (
youtube.com/results?...) are JavaScript-rendered, so WebFetch retrieved only page footers rather than titles, view counts, or channel information. Individual video pages (watch?v=) had the same limitation; titles and channel names were confirmed through the YouTube oEmbed API (youtube.com/oembed). - The oEmbed API does not return views, upload dates, subscriber counts, or comments. Accordingly, view and subscriber counts could not be obtained despite being requested by the playbook. Upload dates are included only where web-search snippets provided them.
- “LTX-2: Open Source AI Video, Released to the Community” (zSvCnVi181A) returned 403 Forbidden during oEmbed access and could not be verified.
- Comments could not be checked for the reasons above; summaries are based on titles and descriptions only.
- Web-search information such as article snippets and related blogs was used to reinforce trend analysis, but every linked YouTube URL was confirmed to exist.
Bluesky
Bluesky — Image and Video Generation AI News (Closed/Open Source)
Accounts
- theinformation.com / techmeme.com / mediagazer.com / lefiltech.fr — RSS-connected news bots spreading OpenAI/Sora scoops in multiple languages.
- ctxwindow.kynth.studio (#TechBluesky) — A watcher that continuously monitors changes to OpenAI’s pricing page.
- gen-ai.news — An individual news account posting Midjourney update alerts.
- oludai.bsky.social — Operates an independent blind human-evaluation leaderboard (olud.ai), ranking open and closed image models hourly.
- aichina.news — A bot-like account reporting nearly daily on Apache-2.0 FLUX LoRAs appearing on Huawei’s Modelers.cn.
- circuitai.bsky.social — A spam/promotional account posting Alibaba WAN 3.0 affiliate links (vidhapi.app) in nearly identical language more than a dozen times per day.
- arxiv-cs-cv.bsky.social — Automatically posts new arXiv papers in image and video generation.
- un1v3rse.bsky.social / carlbetheav.bsky.social / stumblegirl.bsky.social / ttaships.airminded.org — Individual artists posting Midjourney work daily (#midjourney #QP #FoxyFriday).
- rin-yamami.bsky.social / satomin6.alimika.jp — Japanese-language AI illustrators using Qwen-Image-Edit and ComfyUI.
Posts
- theinformation.com (2026-09-10, 👍1 🔁0) — Reported that former OpenAI Sora lead Bill Peebles is launching his own video-generation-AI startup with Hollywood heavyweight Jeffrey Katzenberg (former DreamWorks) and a former Dropbox CFO.
https://bsky.app/profile/theinformation.com/post/3mv6myz7pa52n - ctxwindow.kynth.studio (2026-09-08, 👍0/1 🔁0) — Detected that OpenAI quietly removed sora-2 and sora-2-pro from its pricing page (former prices: 720p $0.05–$0.30), supporting the idea of Sora contraction.
https://bsky.app/profile/ctxwindow.kynth.studio/post/3muxriifygl23 - somewhattolerable.com (2026-09-06, 👍2) — Speculated that “OpenAI discontinued Sora because internal engineers complained that it consumed too many resources.”
https://bsky.app/profile/somewhattolerable.com/post/3muuuazkaf223 - itmatterss.bsky.social (2026-09-11, 👍1) — Reported that Google Gemini became a Windows desktop app launched with Alt+Space, explicitly including Nano Banana for image generation and “Gemini Omni” for video generation.
https://bsky.app/profile/itmatterss.bsky.social/post/3mvayanva5c2n - dimnews.bsky.social (2026-09-11, Russian) — Reported SenseTime’s announcement of the 8-billion-parameter multimodal model “SenseNova-U1.5,” claiming it surpasses Nano-Banana-Pro in visual reasoning.
https://bsky.app/profile/dimnews.bsky.social/post/3mv7u7crnzz25 - gen-ai.news (2026-08-30, 👍1) — Reported that Midjourney V8.2’s editing model was quietly patched following complaints of image-quality degradation.
https://bsky.app/profile/gen-ai.news/post/3mudbgowbnz2s - futureautomationai.bsky.social (2026-09-03, 👍2) — A “Midjourney in 2026” roundup covering V8.2, personalization, moodboards, editing models, and reference features.
https://bsky.app/profile/futureautomationai.bsky.social/post/3mumtnbxuv22d - oludai.bsky.social (2026-09-10/11, 👍1–2) — Its blind-evaluation leaderboard listed Qwen-Image-3.0-Pro as the best downloadable model (#12), only 92 Elo points behind GPT Image 2 (high).
https://bsky.app/profile/oludai.bsky.social/post/3mv5ywkauvt2t - ddigitalmedia.bsky.social (2026-09-10, 👍1) — Argued that “Qwen, Z-Image, and MiniMax are the only players capable of competing with Western LLM/image products.”
https://bsky.app/profile/ddigitalmedia.bsky.social/post/3mv6lyejx3223 - circuitai.bsky.social (2026-09-11–12, repeated posts) — Promoted Alibaba’s video-generation model “WAN 3.0” as a lower-cost, pay-as-you-go competitor to ByteDance Seedance 2.5, supporting clips of up to 30 seconds. The same affiliate copy was posted more than 15 times a day.
https://bsky.app/profile/circuitai.bsky.social/post/3mvbvqsdvr22p - minaxlab.bsky.social (2026-09-11) — Announced the release of ComfyUI v0.35.1, showing continued updates to open-source local-generation tooling.
https://bsky.app/profile/minaxlab.bsky.social/post/3mvasr3l72x2i - aichina.news (2026-09-08–09, repeated posts) — Reported a succession of Apache-2.0 FLUX LoRAs released on Huawei’s Modelers.cn, while repeatedly assessing that many lack trigger words or benchmarks and therefore offer insufficient evidence to claim they surpass Qwen-Image.
https://bsky.app/profile/aichina.news/post/3mv3vq3zdui24 - arxiv-cs-cv.bsky.social (2026-09-11) — Posted a new paper on evaluating and improving “Think-with-Video” reasoning capabilities in video-generation models, “From Evaluation to Enhancement.”
https://bsky.app/profile/arxiv-cs-cv.bsky.social/post/3mvbfk6tunc2j
Signals
- Sora is effectively slowing down: Model removals from the pricing page, the departure of its former lead to establish a competing startup, and speculation that it was cut because of its resource demands all appeared at the same time. OpenAI’s video-generation effort looks defensive.
- Google is expanding in the opposite direction: Nano Banana (image), Nano Banana Pro, Veo, and Gemini Omni (video) are integrated even into a Windows app—an effort to put generative AI into every surface.
- Chinese players lead open source: Qwen-Image-3.0-Pro, Z-Image (Turbo), MiniMax, Alibaba WAN 3.0, and SenseTime SenseNova-U1.5 dominate discussion. Multiple accounts argue that Chinese open models are the only credible challengers to Western image/video AI products.
- The closed-source gap is narrowing: In continuous blind testing by oludai.bsky.social, Qwen-Image-3.0-Pro has closed to within 92 Elo points of GPT Image 2 (high).
- Watch for noise and promotion: Much WAN 3.0 discussion came from circuitai.bsky.social’s repeated affiliate posts and should not be mistaken for organic momentum.
- ComfyUI and LoRA ecosystems continue to evolve: Beyond ComfyUI core updates, FLUX LoRA ports for Huawei Ascend NPUs indicate expansion beyond CUDA hardware.
Limits
- Bluesky’s public APIs (
public.api.bsky.appandapi.bsky.app’sapp.bsky.feed.searchPosts) intermittently returned 403 errors, likely rate limiting. Many queries succeeded only after spacing out retries. - Some keywords—
Runway Gen,Kling AI,Black Forest Labs, andWan 2.5spelling variations—continued to return 403 and were abandoned; substitute searches such asWan videowere used. No direct Runway or Kling posts were included. - The
bsky.app/searchweb UI is a JavaScript SPA and did not render content through fetching, so this work depended entirely on public API JSON. - Literal searches for terms such as “Flux” and “Z-Image” returned mostly unrelated results—Aeon Flux, solar flux, Pokémon Z-A, adult video sites—and required time-consuming filtering.
- The Information’s Katzenberg story and the arXiv paper were verified only at the headline/summary level, not in full text, due to paywall and specialist-paper constraints.
Lemmy
Lemmy — Image and Video Generation AI News
Communities
- !«メールアドレス» — 5,709 total subscribers (893 local). The largest Lemmy hub for image/video-generation tool and model releases. Recent posts are almost entirely daily GitHub/HuggingFace release links from one user,
Even_Adder. - !«メールアドレス» / lemmy.world — Communities for sharing generated artwork rather than news galleries (scores range from -14 to 16).
- !«メールアドレス» — A specialized community for abstract art made with SDXL and custom LoRAs. More than ten posts appeared in September, indicating active participation.
- !ai_reddit@[various instances] — Cross-post aggregation communities for Reddit r/ArtificialInteligence, with little original discussion.
- !«メールアドレス» — General technology news, where Midjourney articles received scores above 70.
- !hackernews@[various instances] — Hacker News cross-posts that surfaced multiple Midjourney stories.
- !techtakes@[various instances] — A strongly AI-critical community where “Midjourney pivots to Theranos” received 40 points.
- !llm@[various instances] — An LLM-specialist community; its GPT-6 Astra release post received a cool response at -3 points.
- !«メールアドレス» — A small community focused on the Perchance text/image-generation tool.
Subscriber counts could not be accurately confirmed outside !«メールアドレス» because APIs returned 404 for some instances.
Posts
Closed-source leaning
- "Midjourney pivots from AI image generation to body scanning medical spa" — !«メールアドレス», 58pt, 2026-06-18. https://www.theregister.com/ai-and-ml/2026/06/18/midjourney-pivots-from-ai-image-generation-to-body-scanning-medical-spa-where-patients-bathe-in-golden-light/5258429 — A startling report that Midjourney pivoted from image generation to a medical-spa business (Midjourney Medical) using ultrasound CT.
- "Midjourney AI pivots to Theranos: Ultrasonic CT" — !techtakes, 40pt, 2026-06-19. https://pivot-to-ai.com/2026/06/19/midjourney-ai-pivots-to-theranos-ultrasonic-ct/ — Commentary sarcastically calling the above “Theranos all over again.”
- "Midjourney Medical" official page — !hackernews, 3pt, 2026-06-18. https://www.midjourney.com/medical
- "Midjourney wants Hollywood studios to reveal the details of their AI usage" — !«メールアドレス», 73pt, 2026-07-05. https://techcrunch.com/2026/07/04/midjourney-wants-hollywood-studios-to-reveal-the-details-of-their-ai-usage/ — Midjourney asks Hollywood studios to disclose how they use AI.
- "Behind the scenes with the Midjourney scanner" (video) — !hackernews, 1pt, 2026-07-06. https://www.youtube.com/watch?v=4nzzpUKhj1M
- "Amnesty International's May 2026 briefing calls leading generative AI systems 'unlawful by design'" — !ai_reddit, 1pt, 2026-06-24. https://www.reddit.com/r/ArtificialInteligence/comments/1ue64s2/ — A human-rights organization criticizes training-data collection by leading closed-source generative AI systems, including image generation, as potentially unlawful.
- GPT-6 Astra release — !llm, -3pt, 2026-09-04. https://openai.com/index/gpt-6-astra/ — The base LLM does not directly generate images/video and instead calls tools through the Responses API. Community response was weak, with a negative score.
Open-source leaning (all !«メールアドレス», posted by Even_Adder)
- MiniMax-H3 video-generation ecosystem tools — More than seven releases between 9/4 and 9/10 (
ComfyUI-MiniMax-H3-W4A4-VSA,MiniMax-H3-Semantic-Bridge,ComfyUI-Ref2VA-VSA,Minimax-h3_Singularity,minimax-h3_fl2v_8Step_motion_enhancer,ComfyUI-VDN-H3,vdn-minimax-h3,FastH3-Ref2V-Stream-Controller). The ecosystem around open-source video model MiniMax-H3 is rapidly expanding through quantization, acceleration, and semantic-bridge tools. Example: https://github.com/sepiablue-ai/ComfyUI-MiniMax-H3-W4A4-VSA - "LLaDA-Image: Building Strong Image Generators" — 3pt, 2026-09-04. https://arxiv.org/abs/2609.03796 — A paper on a 6B-parameter open-source Diffusion Transformer.
- Two LoRAs for Krea 2 Turbo —
F16/krea2-turbo-sda(2pt, 9/10, https://huggingface.co/F16/krea2-turbo-sda) andlvladikov/Krea2-Turbo-Distill-4step-LoRA(2pt, 9/8, claiming a 45% quality improvement through four-step distillation). - "Viggle/Viggle-Animate" — 10pt, 2026-09-05. https://huggingface.co/Viggle/Viggle-Animate — Technology for replacing a character in video using a single redraw frame.
- "How a 20-year-old independent builder from Bihar created an open omni AI model" — !ai_reddit, 0pt, 2026-09-11. https://www.reddit.com/gallery/1wdii5s — An open 5.84B-parameter omni model reportedly surpassed Apple’s AFM 3B on the MATH-500 benchmark.
Signals
- The open-source battleground is shifting from base models to surrounding ecosystems. Around the MiniMax-H3 video model in particular, more than ten releases of quantization tools, acceleration tools, ComfyUI nodes, and LoRAs appeared in one week, concentrating community interest there.
- Lemmy’s coverage of this subject effectively depends on one individual curator (Even_Adder). It resembles a link collection rather than discussion threads, with almost no comments.
- The biggest closed-source story is not a new-model performance race but Midjourney’s pivot to a medical-spa business, a scandalous shift. The tone is more corporate criticism and sarcasm than technology reporting, especially in outlets such as pivot-to-ai.
- Human-rights criticism regarding generative AI legality and data collection, including Amnesty International’s report, also surfaced among closed-source-related stories.
- GPT-6 Astra (released September 2026) does not directly generate images/video, so it has limited relevance to this topic and received a negative reaction on Lemmy.
Limits
- Lemmy’s search API depends on keyword matching. Searches for
Sora,Veo,Runway,Kling,Grok Imagine, andNano Bananaproduced no relevant hits, indicating that these topics are barely discussed on Lemmy. !«メールアドレス»could not load due to HTTP 503, so the dbzer0 instance was used instead.- Community-subscriber APIs (
/api/v3/community?name=...) returned 404 for multiple communities, preventing accurate subscriber counts. - Although 20 relevant posts were collected, Lemmy is a very small platform for this subject, with substantive output concentrated in a few communities and accounts. Original discussions are scarce; most posts are links or reposts.
Pinterest — Image and Video Generation AI News
All 50 pins listed in output/pinterest.pins.md (collected under one uncategorized query, "Image and Video Generation AI News") were opened and visually inspected.
Visual themes
- Neon blue/purple futuristic UI backgrounds dominate. Black backgrounds with circuit-board patterns, glowing AI logos, and holographic humanoid silhouettes recur across infographics and video thumbnails
[1, 3, 4, 5, 34, 37, 39, 41, 48]. - A repeated thumbnail template features a bearded male YouTuber close-up + huge “AI NEWS” lettering + a collage of company logos. Surprised or skeptical faces attract attention alongside Google/OpenAI/Microsoft/Meta/Anthropic logos
[4, 20, 21, 45]. - “Best AI video tool” ranking infographics are extremely common, largely reusing the same template of numbered cards, logos, feature checklists, and “best for” labels. Runway, Pika, Kling AI, Luma (Dream Machine), Canva, Hailuo, and Synthesia make up the standard tool lineup in almost every list
[7, 8, 13, 24, 40]. - Real-vs.-AI detection formats stand out: side-by-side landscape-photo quizzes, deepfake trend explainers, and AI-voice detection meters all center authenticity anxiety
[16, 17, 20, 37]. - Humanoid robots are repeatedly used as generic visual shorthand for AI. Robot news anchors, robots pointing at Earth, and half-cyborg composite faces appear regardless of the actual technical subject
[2, 34, 37, 39, 48]. - One pin displays a material-transformation demo grid, showing the same person, dog, and car video transformed into wooden blocks, origami, Lego, and flowers—one of the most concrete demonstrations of video-model expressiveness
[6]. - Freelancer/service-ad pins are mixed in, including “I make AI videos” ads with WhatsApp contacts and Fiverr listings using generated Jeff Bezos imagery
[15, 32, 42]. - Actual product-news screenshots are a minority. Most pins are mass-produced SEO/affiliate infographics or YouTube thumbnails. Real-news examples include Google Veo 2/3, Microsoft VASA-1, ShengShu Vidu S1, Adobe’s upscaling model, AMD Amuse 3.0, and the ByteDance–MPA copyright agreement
[3, 5, 12, 29, 46, 10, 47, 31].
Notable pins
- #3/#5 — Google announces ultra-high quality video / Veo 2 promo — Official Google promotional materials showing Veo 2 output samples, including a dog swimming underwater, flamingos, and an animated girl in a kitchen.
- #12 — Veo 3 Image-to-Video: Fast Generation & Native Audio via Gemini API — A specific feature update for Veo 3 image-to-video and native audio through the Gemini API.
- #6 — Material-transformation demo grid (Luma-related) — A comparison grid transforming one source video into wooden blocks, origami, Lego, and flowers, making model performance especially tangible.
- #16 — "Why Is Trump Obama AI Video Trending?" — A detailed infographic unpacking a viral AI-video trend, listing tools including Runway ML, Pika Labs, Kaiber, HeyGen, CapCut AI, and DeepBrain AI.
- #29 — ShengShu Technology unveils Vidu S1 — An announcement pin for a Chinese model promoting “real-time, interactive generation.”
- #31 — ByteDance and Hollywood reach global deal on AI copyright — A legal/IP news item claiming ByteDance and the MPA reached an agreement on copyright protections for AI image and video generation tools.
- #44 — "More than 20% of YouTube is now AI-generated" — One of the few pins presenting a concrete statistic.
- #46 — Microsoft VASA-1 — A research demo that generates a talking video from one photo and an audio clip.
- #47 — "Amuse 3.0" (AMD) — An announcement of a new version of an AI image-generation tool.
- #37 — AI-voice detection infographic ("97% likely AI generated") — An educational pin focused on authenticity assessment and AI ethics.
Signals
- The center of discussion remains closed-source commercial tools—Runway, Pika, Kling AI, Luma, Canva, Hailuo, Synthesia, and D-ID. Their logos appear in nearly every ranking infographic. No pins foregrounding open-source/self-hosted workflows such as Stable Video Diffusion, ComfyUI workflows, or Open-Sora were found.
- Google Veo 2/3 is the most visible official product news, emphasizing image-to-video and native audio through the Gemini API.
- Among Chinese players, ShengShu’s Vidu S1 appears through its “real-time, interactive generation” framing.
- Deepfakes and authenticity anxiety—the Trump-Obama video trend, AI-voice detection, and “which is AI?” quizzes—form a consistent secondary theme on Pinterest.
- Industry-structure and legal news also appears, including AI-generated-content share on YouTube (more than 20%) and the ByteDance–MPA copyright agreement.
- Technical depth—model architectures, training methods, benchmark comparisons—is almost absent. Pinterest naturally skews toward tool introductions, how-tos, and explanations of viral phenomena.
Limits
- The collection worker used only one query (
"Image and Video Generation AI News") without categorizing by genre such as open versus closed source. This analysis therefore covers all 50 pins together; genre-level segmentation was left to the later integration stage. - Some pins were only weakly related to the topic or were generic stock imagery, including
#35(world-map current-affairs image),#43(a statue-like cello player), and#22(Adobe Stock icon material). Pinterest matching appears loose, creating noise. - Pinterest does not display post dates, so individual pin dates cannot be determined except from labels such as “2025” or “2026” in titles.
- Pin titles are cached Pinterest text and may differ from the actual date and details of the linked article. This page verifies only the image and title-level information.
Recommended actions
- Re-run X research using specific names such as Midjourney, Sora, Runway, Kling, Flux, and ComfyUI rather than trending terms.
- Monitor the Elo-score gap between Qwen-Image-3.0-Pro and GPT Image 2 to track the open/closed performance gap.
- Continue following Bluesky pricing-page-monitoring accounts to verify the claim that Sora is slowing down.
- Broaden Reddit search terms to reach the target of 20 threads.
- Do not treat Pinterest ranking infographics or Bluesky affiliate-post floods as reliable on their own; use them only after corroboration elsewhere.
Collected images


















































Data quality notes
X had zero image/video-generation-AI posts because of poor search-term selection. Reddit had only 12 posts, short of the target of roughly 20. YouTube’s JavaScript-rendering limitations prevented retrieval of view counts and subscriber counts, while Bluesky collection was abandoned for some keywords due to repeated 403 errors.



