KEN’S CAT LOG
Today's LLM News

Daily LLM News — 2026-09-12

Today's biggest story was not a model announcement, but Anthropic's threat intelligence report (Houthi weapons development, Iranian warship surveillance, and Russia-linked hackers abusing Claude in attacks on Ukraine), which independently drew top-tier engagement on X and Lemmy.

Daily LLM News — 2026-09-12

The most-discussed topic today was not new-model announcements themselves, but AI misuse and trust. On X and Lemmy, Anthropic's threat intelligence report—covering Houthi weapons development using Claude Code, Iranian monitoring of U.S. naval vessels, and Russia-linked hackers stealing Ukrainian military technology—independently emerged as the day's biggest story. Meanwhile, Reddit and Bluesky both saw simultaneous discussion of allegations that OpenAI may have "stolen" mathematicians Buckmaster and Alpoge's Navier–Stokes-related work from ChatGPT session histories. Elsewhere, YouTube, Bluesky, and Lemmy saw a succession of new models in the spotlight, including Claude Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash, DeepSeek V4.1 Flash, GLM-5.3, and Sakana AI's Fugu Max/Ultra v2, with a consistent closed-models-versus-open-weights divide. Overall, the main themes are clear, but some platforms did not meet the collection-completeness threshold because of sampling and search constraints, so confidence remains moderate.

Across platforms

Platform by platform

Reddit (first in collection order): Twelve threads were collected from nine subreddits using the single search term "Daily LLM News," but no threads reporting new-model announcements or price changes themselves were found; discussion, sarcasm, and reactions to existing news dominated. The main flashpoints were the OpenAI Navier–Stokes "theft" allegation and backlash against the WSJ's open-weights critique. The coverage window was also spread across five days, September 6–10, 2026, rather than tightly focused on "today" as required by the brief.

X (second): Because the session's trend list was based on a connection identified as Croatia, virtually the only LLM-related search term among ten collected trends was "Anthropic." Just six of ten posts directly related to the LLM brief, falling short of the threshold of ten. Among those six, Anthropic's threat intelligence report about the Houthis and Iran had by far the greatest reach, while its economic-scenarios report and Claude Code open-sourcing received a contrastingly positive reception.

YouTube: The ten-item completion threshold was met, with near-comprehensive coverage of this week's major model announcements: Claude Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash, DeepSeek V4.1 Flash, GLM-5.3, Muse Spark 1.3, and Sakana Fugu Ultra. However, JavaScript-rendering limits prevented collection of exact view and subscriber counts, and many dates were inferred from relative timestamps.

Bluesky: Because the official search API returned HTTP 403, twelve items were collected through a workaround: tracing feeds from accounts found via web search, including Ethan Mollick, TestingCatalog, and Sakana AI. Likes and repost counts were confirmed for ten items. Sakana AI's Fugu Max/Ultra v2 launch made the open-weights position most explicit, but the collection was skewed toward accounts discovered incidentally.

Lemmy: Twelve items were collected across communities including !technology, !localllama, and !fosai. Closed-model discussion was concentrated almost entirely in the security-and-misuse context of Anthropic's threat report, while open-weights discussion centered on practical reports such as Qwen quantization and K2 Horizon in LocalLLaMA. Very few posts were dated today (09-12); recent posts were spread from 09-04 through 11.

All five platforms explicitly named in the brief were investigated; the absence of other platforms is not a gap.

What to watch

Recommendations

  • Review Anthropic's original threat intelligence report and verify the primary sources behind its claims about weapons misuse and surveillance use.
  • Replace X search terms based on connection-local trends with a fixed list explicitly naming companies such as OpenAI, Google, Meta, xAI, and Mistral.
  • Supplement Reddit collection beyond "Daily LLM News" with model and company names so that new-model announcements themselves can be captured.
  • Continue tracking Sakana AI Fugu Max/Ultra v2 and OpenAI Astra in future collection cycles.
  • Check next time whether Bluesky's searchPosts API HTTP 403 issue has been resolved; return to direct search if it has.
  • Prioritize checking whether OpenAI has issued an official statement on the Navier–Stokes allegations.

Data quality

X trend selection was not optimized for the LLM brief, and only six of the required ten items were LLM-related. Reddit and Lemmy met their item-count thresholds, but their coverage windows were spread over periods broader than "today"—five days for Reddit and 09-04 through 11 for Lemmy—so they are more a collection of discussion around existing news than news unique to the day. Bluesky's official search API was unavailable due to HTTP 403, so its account-tracing method created collection bias. YouTube met the ten-item threshold, but JavaScript rendering prevented collection of quantitative data such as view and subscriber counts.

Platform summaries

Reddit

Reddit — Daily LLM News

Where

Twelve threads were collected from nine subreddits using the search term "Daily LLM News."

Subreddit Members Collected threads
r/NoStupidQuestions 7,480,942 3
r/singularity 3,970,329 2
r/ArtificialInteligence 1,924,916 1
r/artificial 1,336,553 1
r/LocalLLaMA 821,987 1
r/WritingWithAI 165,025 1
r/SillyTavernAI 128,146 1
r/BetterOffline 55,091 1
r/Maxcactus_TrailGuide 5,205 1

r/NoStupidQuestions and r/singularity had the most items, but their content was mainly general discussion along the lines of "what are LLMs?" By contrast, one item (#10) from tiny r/Maxcactus_TrailGuide, with just 5,205 members, received an unusually viral 2,213 points.

What people say
  • Claims related to OpenAI's Millennium Prize problem (Navier–Stokes), along with the associated "idea theft" allegation, were today's largest flashpoint. Thread #5 (r/singularity, 209pt/69 comments, 2026-09-10, https://www.reddit.com/r/singularity/comments/1wcnyay/) reported that OpenAI was approaching a solution to the Navier-Stokes equations, while OpenAI denied allegations that Codex-use histories from mathematicians including Tristan Buckmaster (referenced in a NYT article) may have influenced the model. u/FateOfMuffins (61pt) commented that this might suggest the most recent major training data cutoff was July 3.
  • Reactions to the same story split by subreddit. In r/singularity's thread #6, "LLMs will never be smarter than Redditors" (271pt/159 comments, 2026-09-09, https://www.reddit.com/r/singularity/comments/1wbnhgt/), u/oilybolognese (201pt) mocked skeptics. In r/SillyTavernAI's thread #7 (68pt/57 comments, 2026-09-10, https://www.reddit.com/r/SillyTavernAI/comments/1wclha4/), the conclusion was the reverse. u/perthro_anon (140pt) sarcastically said they were worried OpenAI would steal the smut they wrote and take credit for it, leading into arguments in favor of local LLMs.
  • A WSJ article criticizing open weights and targeting closed-model competitors drew strong backlash in r/LocalLLaMA. Thread #12 (513pt/236 comments, 2026-09-08, https://www.reddit.com/r/LocalLLaMA/comments/1wa9309/) covered a WSJ opinion piece claiming that unregulated open-weight AI is an invitation to disaster, including a claim that it answered a question about synthesizing poliovirus. u/3169676 (557pt, the thread's top-scoring comment) pushed back: should only wealthy billionaires be able to ask questions of "regulated" AI models? u/inotparanoid called it a hit piece intended to protect Anthropic and OpenAI's post-IPO share prices.
  • Distrust of stories about "defectors" raising AI-safety concerns appears in two threads. Thread #4 (r/ArtificialInteligence, 0pt/32 comments, 2026-09-09, https://www.reddit.com/r/ArtificialInteligence/comments/1wbu9c6/) discussed Anthropic researcher Jacob Coxon leaving the industry over concern about a race to develop uncontrollable systems, but u/MiloGoesToTheFatFarm dismissed him as someone who had merely done low-level work like data entry, not model design. Thread #9 (r/BetterOffline, 299pt/167 comments, 2026-09-10, https://www.reddit.com/r/BetterOffline/comments/1wcxjv0/) discussed another AI doomer featured on CNN; u/Smurfette2016 (208pt) said that someone with six weeks of employment and a Twitter account less than a month old with zero followers getting 130 million views on one post was obviously orchestrated.
  • Pushback against the "it's only next-token prediction" line recurred across several threads, though the mood today was somewhat tired and repetitive. In thread #1 (r/NoStupidQuestions, 1,285pt/358 comments, 2026-09-10, https://www.reddit.com/r/NoStupidQuestions/comments/1wcsbr2/), u/skmchosen1 (16pt, self-described AI researcher) explained that next-token prediction is the pretraining objective, but post-training introduces other objectives, such as rewarding successful solutions to math problems. u/danderzei (62pt) reopened the same kind of discussion in thread #3 (r/artificial, 67pt/221 comments, 2026-09-06, https://www.reddit.com/r/artificial/comments/1w8m8rw/).
  • A scientist's claim that LLMs are cognitive viruses for humans saw exceptional growth despite originating in a small subreddit. In thread #10 (r/Maxcactus_TrailGuide, 5,205 members, 2,213pt/180 comments, 2026-09-09, https://www.reddit.com/r/Maxcactus_TrailGuide/comments/1wbi2d7/), u/MF_D000M (58pt) commented that LLMs are accelerating humanity toward the final stage of mob rule. u/TurgorFervor compared religion to a virus as well.
  • Working users also voiced practical frustrations with writing quality on free tiers. In thread #8 (r/WritingWithAI, 18pt/35 comments, 2026-09-10, https://www.reddit.com/r/WritingWithAI/comments/1wcunsz/), the poster said GPT Sol is good for research but writes robotically, Claude Sonnet's free-tier restrictions are severe, and Gemini is unintelligent. u/whitemisandry (4pt) said they use Claude Opus 4.8 but provide it with a long instruction system before writing.
  • For the basic question "why are LLMs getting smarter?", an explanation based on overall system architecture won support. In thread #11 (r/NoStupidQuestions, 1pt/9 comments, 2026-09-09, https://www.reddit.com/r/NoStupidQuestions/comments/1wbpr47/), u/Jolly-Rip5973 (2pt) explained that standalone LLMs are nearing a plateau and that current progress comes from system-level improvements such as harnesses, tool calls, and coordination among multiple models.
Signals
  • Rising sentiment: distrust and backlash toward closed-model companies, especially OpenAI. More than the Millennium-Prize-scale result itself, the allegation that it stole ideas (#5, #7) and backlash to the WSJ article criticizing open weights (#12) generated Reddit's heat today. Open weights are defended as symbols of free expression and independence.
  • Dismissed sentiment: Posts by AI-safety "defectors" (#4, #9) received high engagement, but most commenters dismissed their allegations as publicity, PR, or suspicious. Fatigue and cynicism toward safety rhetoric itself are strong.
  • Surprise: A very small subreddit with only 5,205 members (#10) recorded 2,213 points, surpassing far larger communities such as r/singularity with 3,970,329 members. It likely went viral through r/all or similar channels, breaking the usual relationship between community size and score.
  • Conflict structure: On the same OpenAI Codex/Millennium Prize news, r/singularity (#5, #6) leaned toward "OpenAI is right and the skeptics are ridiculous," while r/SillyTavernAI (#7) leaned toward "OpenAI stole it; local LLMs are necessary"—the same story was interpreted in completely opposite ways by different communities.
Limits
  • Only the single search term "Daily LLM News" was used; no specific model or company names were searched. Therefore, the twelve collected threads focused on discussion, reactions, and sarcasm around existing news rather than primary posts about newly announced models, price cuts, or benchmarks. No threads directly reporting new-model releases or API price changes were found.
  • Three r/NoStupidQuestions items (#1, #2, #11) were beginner questions about how LLMs work in general. They are enduring discussions rather than today's news, with little day-specific material beyond recency.
  • At most six top comments per thread were available, so the broader tone of the comments could not be verified. For example, only six of 358 comments were available for #1 and six of 236 for #12.
  • The coverage period was spread across the five days from 2026-09-06 to 2026-09-10, rather than posts strictly limited to today (2026-09-12). The worker's collection window may have been broader than the brief intended.

X

X — Few LLM-related topics; the central story is Anthropic's report that Claude was misused for weapons development

The session's Explore page was displayed as access from Croatia, as indicated by items labeled "Trending in Croatia." The top ten trends were "Yemen," "Saudi," "$song," "Charlie Kirk," "Britain," "Anthropic," "Apple," "Sony," "one ai os," and "Chelsea." This was the Explore ordering for Croatia—or at least this session's connection origin—not a universal global trend list. Collection was conducted by searching these ten terms. The only substantive LLM-related material came from posts found under "Anthropic," plus Anthropic threat-intelligence posts incidentally found through the "Yemen" search.

Accounts

All 40 collected accounts contributed only one post each; no account appeared more than once. The six accounts posting LLM-related content were:

Account Content
@tleilax___(Yet another commodity guy) Said a Yemen-based group used Claude Code to develop guided rockets, ballistic missiles, and a hypersonic glide vehicle. 5,022 likes
@IranTimes72(Ebrahim Zolfaghari) Reported that Iran used Claude AI to monitor U.S. warships, noting that Anthropic also identified a Yemen-based group
@BrianRoemmele A long post mocking a person who spent "six weeks" at Anthropic and said AI would destroy humanity within ten years
@restitutorII(Restitutor Orientis) Introduced Anthropic's report depicting three U.S. economic scenarios for 2030, including GDP +1.6%
@walterkirn A short sarcastic post about the same person, saying mainstream media elevated them like Dr. Fauci
@claudecode84(Claude code research lab, Japanese-language account) Reported that Anthropic's top executive open-sourced their entire Claude Code environment. About 724,000 views, the highest among LLM-related posts
Posts
  1. @tleilax___, 2026-09-10 (5,022 likes, 359 reposts, 21 replies, about 264,000 views; found through "Yemen" search)

    "One Yemen-based group used multiple Claude Code instances while working on guided rockets, ballistic missiles and a hypersonic-glide project."
    A post apparently quoting Anthropic's threat intelligence report. It says a Yemen-based group used multiple Claude Code instances while working on weapons development, and was the most widely shared LLM-related post.

  2. @IranTimes72, 2026-09-11 (238 likes, 34 reposts, 8 replies, about 15,000 views; found through "Yemen" search)

    Iran used Claude AI to monitor U.S. warships in the Middle East and identify potential communications vulnerabilities. ... Anthropic also identified a Yemen-based [group]
    Another post about the same Anthropic report. It reported that Iran combined warship tracking, military imagery, and commercial satellite images and analyzed them with Claude; it is spreading alongside post 1.

  3. @restitutorII, 2026-09-10 (458 likes, 69 reposts, 41 replies, about 57,000 views; found through "Anthropic" search)

    Anthropic in its report imagines 3 possible futures for the American economy in 2030 based on the power and adoption of AI: 1- Modest scenario: AI has an impact comparable to the Internet. GDP +1.6%, few disruptions to employment. 2- Substantial scenario: AI can perform about...
    An introduction to Anthropic's economic-scenarios report, which forecasts AI's impact on the U.S. economy in 2030 at three levels. Separate from the weapons-misuse story, Anthropic's economic-impact analysis was also discussed that day.

  4. @claudecode84, 2026-09-10 (2,522 likes, 263 reposts, 24 replies, about 724,000 views; found through "Anthropic" search)

    Anthropic最高責任者本人が自分の「Claude Code環境」を丸ごとオープンソース化しています
    A Japanese-language account's post. By views alone, it was the largest LLM-related spread. It says an Anthropic leader released their own Claude Code environment, although the post text itself did not provide a source link or further details.

  5. @BrianRoemmele, 2026-09-10 (790 likes, 187 reposts, 105 replies, about 99,000 views; found through "Anthropic" search)

    TEMU HARRY POTTER SAYS "AI WILL END ALL HUMAN LIFE IN 10 YEARS!" ... His SIX WEEKS at Anthropic...
    A post mocking someone who spent only six weeks at Anthropic and said AI would destroy humanity in ten years. Its 105 replies and high engagement make skepticism and backlash toward AI-safety rhetoric visible.

  6. @walterkirn, 2026-09-10 (722 likes, 91 reposts, 39 replies, about 28,000 views; found through "Anthropic" search)

    The mainstream media boosted this dude like he was, well, Dr Fauci or something. I implore you to see clearly how this works. And who works it.
    A post apparently referring to the same person as #5. Though unnamed, its context, search term, and date align; it sarcastically criticizes mainstream media for elevating the individual.

Signals
  • The most-discussed issue is Claude's military and intelligence use. A Yemen-based group's weapons development (1) and Iran's warship surveillance (2) both appear to be based on Anthropic's threat intelligence report, and were incidentally found through the geopolitical term "Yemen." This was the most widely shared LLM-related story in the session.
  • Anthropic's economic report (3) and the Claude Code open-sourcing story (4) had positive, informational tones, sharply contrasting with misuse-related news.
  • Skepticism and mockery dominate responses to claims that AI will destroy humanity within ten years (5, 6). Backlash against media promotion of a former Anthropic employee whose tenure was only six weeks appears stronger than support.
  • Surprise: No posts on open-weight players such as Llama, Qwen, DeepSeek, or Mistral appeared in this collection. No concrete posts about benchmarks or API price changes were found either.
  • Of the 40 collected posts and ten search terms, only six directly related to the LLM brief. The remainder concerned general Yemen-Saudi conflict news, political controversy around the first anniversary of Charlie Kirk's assassination, Britain-Israel diplomacy, reactions to an Apple event, Sony and gaming topics, and Chelsea football memes.
Limits
  • Explore, the trend list, was displayed according to Croatia, and does not represent global trends. Because it depends on the session's connection origin, rankings would likely differ in other countries or regions.
  • Search-term selection was not optimized for the LLM brief. The ten terms used—"Yemen," "Saudi," "$song," "Charlie Kirk," "Britain," "Anthropic," "Apple," "Sony," "one ai os," and "Chelsea"—were copied directly from Explore trends, and only "Anthropic" was LLM-related. No searches targeted developments at OpenAI, Google/Gemini, Meta AI, xAI/Grok, Mistral, or other companies, so coverage was not comprehensive for either closed-model or open-weight players.
  • The ten-item completion threshold was not reached: only six LLM-related items could be written up out of 40 posts. This collection included no posts concerning new-model launches, pricing changes, benchmarks, or concrete incident cases.
  • The session itself operated normally; there was no expired login or page-fetch failure. The missing coverage was caused by search-term selection, not session malfunction.

YouTube

YouTube — Daily LLM News (September 12, 2026)

Channels
  • Anthropic (official channel) — Posted an announcement video for Claude Fable 5.1.
  • ByteForward — A channel focused on comparisons between closed models and open weights.
  • WorldofAI (@intheworldofai) — Frequently posts practical new-model tests and AI-news roundups.
  • AI WITH Rithesh — Covers new-model benchmarks and pricing.
  • Prompt Engineering (@engineerprompt) — Covers prompt design and new models.
  • AI Coding Daily — Provides hands-on coding-benchmark reviews.
  • RandomAI — Covers leaks and rumor-based analysis.
  • TheAIGRID — Provides rapid explainers for new-model launches.
  • AI時短ラボ — A Japanese-language breaking-AI-news channel.

Because channel pages use JavaScript rendering, subscriber counts could not be confirmed directly and exact figures were unavailable (see Limits).

Videos
  1. Introducing Claude Fable 5.1 — Anthropic — September 1, 2026 — https://www.youtube.com/watch?v=ROF2Nv_KjOM
    Anthropic's official launch video. It positions Claude Fable 5.1 as the latest version of its highest-performance class and says its Terminal-Bench-Science score rose from 24.7 to 52.6 from version 5 to 5.1.

  2. GPT-6 Astra vs DeepSeek V4.1 Flash… Is It Even Close? — ByteForward — around September 11, 2026 (posted one day earlier) — https://www.youtube.com/watch?v=YnFKMybL7gs
    A hands-on comparison between closed GPT-6 Astra and open-weight DeepSeek V4.1 Flash. It presents a benchmark claiming V4.1 Flash reached 98% of Astra's score on design tasks at 1.4% of the cost.

  3. DeepSeek V4.1 Flash Is INSANELY GOOD! Fast, Cheap, Powerful! (Fully Tested) — WorldofAI — early September 2026 — https://www.youtube.com/watch?v=T2dnchLabZQ
    Tests coding, 3D simulation, agent behavior, and image input, praising speeds of up to 427 tokens per second and low cost.

  4. GLM-5.3 Released: Everything You Need to Know About Z.ai's New Coding Model (Benchmarks, Price) — AI WITH Rithesh — mid-August 2026 — https://www.youtube.com/watch?v=4RFo47KJ4q4
    Explains benchmarks and pricing for Z.ai's open-weight GLM-5.3. It emphasizes performance gains achieved solely through later-stage training while retaining the same 743B base model.

  5. 【速報】GLM-5.3登場 コーディングとサイバー能力で最強クラス 期待のモデル!! — AI時短ラボ — mid-August 2026 — https://www.youtube.com/watch?v=Ca0PHSjrU8M
    A Japanese-language breaking-news video about GLM-5.3. It describes the model as among the strongest for coding and cybersecurity-related tasks.

  6. Gemini 3.8 Flash: The model no one expected! — Prompt Engineering — around September 2–3, 2026 — https://www.youtube.com/watch?v=UvrAYDgobSw
    Calls Google's Gemini 3.8 Flash, introduced on September 2, 2026, a model nobody expected, claiming its free tier approaches Claude Opus 5 performance.

  7. I Tested NEW Muse Spark 1.3 on Coding: Meta Joins Frontier LLMs? — AI Coding Daily — early September 2026 (posted one week earlier) — https://www.youtube.com/watch?v=lxljOqB1YUI
    Tests Meta's planned open-weight model Muse Spark 1.3 on coding tasks and considers whether Meta has caught up to the frontier tier.

  8. OpenAI Astra Just Got Leaked and It's INSANE (Full Breakdown) — RandomAI — late August 2026 (posted two weeks earlier) — https://www.youtube.com/watch?v=2sh0_oV_r04
    A breakdown of leaked information about OpenAI's upcoming Astra model. It should be treated cautiously because it is rumor-based and predates a formal announcement.

  9. DeepSeek V4.1 GA Soon, GPT-5.6 SOL Nerfed? HUGE Fable Update, US AI BAN Protests, & More! AI NEWS — WorldofAI — early September 2026 — https://www.youtube.com/watch?v=LgGJz528Boo
    A weekly AI-news roundup covering the expected general release of DeepSeek V4.1, Fable updates, protests against U.S. AI regulation, and more.

  10. Sakana AI's New "Fugu Ultra" Beats Claude Fable 5 (Sakana Fugu) — TheAIGRID — late June 2026 — https://www.youtube.com/watch?v=FwA1bcpSGiM
    Introduces the original Fugu Ultra from Japanese company Sakana AI, a multi-agent-oriented model series, and examines its claim to outperform Claude Fable 5. Dedicated videos for its successors Fugu Ultra v2 / Fugu Max, released September 11, had not yet been found (see Limits).

Signals
  • Closed-model camp: Anthropic (Fable 5.1, 9/1), Google (Gemini 3.8 Flash, 9/2), and OpenAI (Astra, unannounced and in the leak stage) are moving in rapid succession, and YouTube review videos were concentrated within days of announcements.
  • Open-weight camp: DeepSeek (V4.1 Flash), Z.ai (GLM-5.3 series), Meta (Muse Spark 1.3), and Sakana AI (Fugu Max / Fugu Ultra v2) are framed as getting close to closed models at lower cost. In particular, the comparison claiming DeepSeek V4.1 Flash achieves 98% of Astra's score at 1.4% of the cost is frequently cited among creators.
  • Japanese-language channels such as AI時短ラボ also cover Chinese open-weight models such as GLM-5.3 quickly; the time lag is only slightly longer than in English-language coverage, generally within one to two weeks of release.
  • Sakana AI's Fugu Max / Fugu Ultra v2, released September 11, 2026, has been covered by news sites and blogs, but dedicated YouTube explainers had not yet been found at the time of this research; only recent videos on the original Fugu Ultra were available.
Limits
  • Because YouTube search-result pages (youtube.com/results?...) and video pages (youtube.com/watch?...) are JavaScript-rendered, WebFetch returned only footers rather than body content such as titles, channel names, views, and dates. Information was supplemented using web-search snippets and noembed.com, which returns only titles and channel names.
  • For that reason, exact view counts and subscriber counts could not be obtained. Many dates were inferred from relative wording such as "X days ago" or "X weeks ago," and the exact upload date could not be identified for some videos.
  • Dedicated YouTube videos about Sakana AI's newest Fugu Max / Fugu Ultra v2 models, released 9/11, could not yet be found. Only links embedded in news articles and playlists were found; whether a video itself exists remains unconfirmed.
  • The ten-item threshold was met, but one item, Sakana Fugu Ultra, covers the original model rather than the latest v2 announcement itself.

Bluesky

Bluesky — Model launches flow through as "bot-like breaking news," while the Navier–Stokes controversy and Sakana AI's open-weights declaration stand out

Note on collection method

Bluesky's official full-text search API (public.api.bsky.app/xrpc/app.bsky.feed.searchPosts) and web search (bsky.app/search) could not be used from this session's fetch tool: the former returned HTTP 403, while the latter returned only empty content because JavaScript was not executed (see "Limits"). The workaround was to identify relevant bsky.app posts and accounts via web search, then read app.bsky.feed.getAuthorFeed for the discovered handles—which did return HTTP 200—to confirm post text, dates, likes, and reposts.

Accounts
Account Profile Link
Ethan Mollick (@emollick.bsky.social) Wharton professor who posts long, frequent threads about the practical impact of frontier models https://bsky.app/profile/emollick.bsky.social
🚨 AI News | TestingCatalog (@testingcatalog.com) Breaking-AI-news account that posts one-line updates on new models and features https://bsky.app/profile/testingcatalog.com
Sakana AI (@sakanaai.bsky.social) Tokyo-based research lab focused on open weights and multi-agent orchestration https://bsky.app/profile/sakanaai.bsky.social
hardmaru / David Ha (@hardmaru.bsky.social) Sakana AI cofounder and author of the official announcement thread https://bsky.app/profile/hardmaru.bsky.social
Quantian (@quantian.bsky.social) Individual account receiving attention for a sarcastic post about the OpenAI Navier–Stokes allegations https://bsky.app/profile/quantian.bsky.social
Zach Weinersmith (@zachweinersmith.bsky.social) Science-comics author who explained the controversy in a thread https://bsky.app/profile/zachweinersmith.bsky.social
ardalis / Steve Smith (@ardalis.com) Developer who posted a short joke mocking Claude Opus https://bsky.app/profile/ardalis.com
Posts
  1. Sakana AI (@sakanaai.bsky.social) / 2026-09-11T20:43Z / 9 likes, 1 repost
    https://bsky.app/profile/sakanaai.bsky.social/post/3mvbglsna4k25

    "Fugu Max is live on Open Router! 🐡 Our learned multi-agent orchestration engine routes tasks across open-weights and specialized models for Pareto-frontier efficiency."
    Announcement that Fugu Max is available on OpenRouter. It explicitly presents an orchestration approach that combines open-weight and specialized models for optimal task routing.

  2. hardmaru (@hardmaru.bsky.social, Sakana AI cofounder) / 2026-09-11T02:57Z / 48 likes, 5 reposts
    https://bsky.app/profile/hardmaru.bsky.social/post/3mv7kz2uplk2u

    "Introducing Fugu Max and Fugu Ultra v2: Orchestrating the Pareto Frontier. The AI industry has spent a decade optimizing along a single axis: build bigger, more expensive models."
    A follow-up post explicitly states that "Fugu Ultra v2 does all of this without Fable 5, Fable 5.1, or GPT-6-Astra in its agent pool," presenting independence from Anthropic and OpenAI frontier models as a strength. This was the clearest open-weights position found today.

  3. TestingCatalog (@testingcatalog.com) / 2026-09-11T07:38Z / 1 like, 1 repost
    https://bsky.app/profile/testingcatalog.com/post/3mva2p4ykhd2s

    "OpenAI launches ChatGPT for Financial Services"
    Breaking news that ChatGPT incorporating GPT-6 Astra is available to financial institutions for research, modeling, and document preparation.

  4. TestingCatalog (@testingcatalog.com) / 2026-09-10T20:45Z / 0 likes, 0 reposts
    https://bsky.app/profile/testingcatalog.com/post/3mv6waz3vpu2t

    "OpenAI launches GPT-Live-1 for full-duplex voice agents"
    Breaking news about a new full-duplex voice model that can listen and speak simultaneously.

  5. TestingCatalog (@testingcatalog.com) / 2026-09-03T09:19Z / 2 likes, 0 reposts
    https://bsky.app/profile/testingcatalog.com/post/3mum4nfbitu2s

    "Google releases Gemini 3.8 Flash and Flash Cyber"
    Announcement of Gemini 3.8 Flash for coding, agents, and multi-step reasoning, alongside Flash Cyber, a variant focused on finding and fixing vulnerabilities and restricted to trusted defensive users.

  6. TestingCatalog (@testingcatalog.com) / 2026-09-02T00:00Z / 2 likes, 0 reposts
    https://bsky.app/profile/testingcatalog.com/post/3muimw7yugw2q

    "Anthropic launches Claude Fable 5.1 and Mythos 5.1"
    Simultaneous launch of the general-purpose Fable 5.1 and Mythos 5.1, restricted to "vetted defenders and scientists," with claims of lower cache costs and better performance.

  7. Ethan Mollick (@emollick.bsky.social) / 2026-09-11T17:38Z / 40 likes, 6 reposts
    https://bsky.app/profile/emollick.bsky.social/post/3mvb4b7idf22v

    "Here's a automated forecasting system to estimate catastrophic risk from some major experts on forecasting. The models predicts the chance of an AI-generated mass catastrophe as 0.47% by 2030..."
    An introduction to a system that automatically estimates the probability of AI-driven mass catastrophe using expert forecasting methods. A follow-up reply notes that the figure assumes no policy intervention.

  8. Ethan Mollick (@emollick.bsky.social) / 2026-09-11T04:59Z / 65 likes, 2 reposts
    https://bsky.app/profile/emollick.bsky.social/post/3mv7rswik2s2v

    "Astra's turn: 'Make a game about Imminence...'"
    A game-making demo created with GPT-6 Astra. It is framed as a comparison of creative ability with a Claude Fable 5.1 game posted the previous day.

  9. Ethan Mollick (@emollick.bsky.social) / 2026-09-10T17:32Z / 117 likes, 18 reposts
    https://bsky.app/profile/emollick.bsky.social/post/3mv6lhq5lyc2v

    "What you are seeing in math right now is a consequence of the [frontier models' capability jump]..."
    Part of a thread on mathematicians' use of AI and what it means for AI to discover superhuman proofs. It is read in the same context as the Navier–Stokes controversy that spread on Reddit's r/singularity, and was the highest-liked item in this collection.

  10. Quantian (@quantian.bsky.social) / 2026-09-08T07:13Z
    https://bsky.app/profile/quantian.bsky.social/post/3muyhwbcdys2x

    "Dancing very very VERY close to saying 'OpenAI stole our almost-complete proof of Navier-Stokes from internal access to our ChatGPT sessions because they were desperate to scoop Anthropic' which is absolutely 100% a believable thing OpenAI would do right now"
    A post discussing allegations that OpenAI stole mathematicians Buckmaster and Alpoge's Millennium-Prize-scale Navier–Stokes work from ChatGPT session histories—the same topic as the Reddit r/singularity thread from that day.

  11. Zach Weinersmith (@zachweinersmith.bsky.social) / 2026-09-08T09:45Z
    https://bsky.app/profile/zachweinersmith.bsky.social/post/3muyqfsvt6k2f

    "For anyone not following the drama, there's been apparently a big development in ai math, along with some accusations of very bad behavior by OpenAI. ... Buckmaster and Levent Alpoge made a big discovery on Navier-Stokes."
    The opening of an explanatory thread summarizing the course of the controversy above.

  12. ardalis / Steve Smith (@ardalis.com) / 2026-09-10T12:35Z
    https://bsky.app/profile/ardalis.com/post/3mv62uanbkk2b

    "LOL #ai #claude opusfived.dev"
    A short joke post mocking something related to Claude Opus. Details are unclear, but the context and linked domain name, "opusfived," suggest a meme about a new Opus version.

Signals
  • LLM discourse on Bluesky is centered on practitioner analysis and breaking-news bots; no controversy or geopolitical story comparable to Claude's weapons use on X was found in this collection. Accounts such as TestingCatalog list new models and features matter-of-factly, while practitioner accounts such as Ethan Mollick discuss usage and societal effects—a two-layer structure.
  • Sakana AI made the open-weights position explicit today. hardmaru, Sakana AI's cofounder, stated that Fugu Ultra v2 depends on none of Fable 5, Fable 5.1, or GPT-6-Astra, selling multi-agent orchestration that avoids lock-in to frontier closed models. This is a concrete example of the closed-versus-open-weights divide that did not appear in the X or Reddit collection.
  • The Navier–Stokes/OpenAI allegations were discussed around the same time on both Bluesky and Reddit. Math and AI commentators such as Quantian and Zach Weinersmith picked up the same issue across platforms, making it the best-corroborated cross-platform story.
  • Engagement is more than an order of magnitude smaller than on X. The largest like count confirmed here was Ethan Mollick's 117 (#9), very different in scale from X posts with thousands of likes. For this topic, Bluesky appears to be a technical and observer-oriented channel focused on quality over volume.
Limits
  • The official search API (app.bsky.feed.searchPosts) consistently returned HTTP 403 from this session's fetch tool, so keyword-based full-text search could not be performed. app.bsky.actor.getProfile and app.bsky.feed.getAuthorFeed returned 200 from the same domain, suggesting only the search endpoint was blocked. The bsky.app/search web interface and embed pages also returned only pre-JavaScript HTML, so post lists could not be read.
  • Because of these constraints, collection relied on the indirect method of finding relevant bsky.app links via web search and tracing discovered accounts' feeds. This was not platform-wide keyword coverage and is biased toward accounts incidentally found by web search, including Mollick, TestingCatalog, and Sakana AI. Accounts regularly covering individual Chinese or Meta open-weight models such as DeepSeek, Qwen, Llama, and Mistral were not found through this method.
  • No items posted today, 2026-09-12, were found. The newest posts were from 2026-09-11, possibly due to search-engine indexing delays.
  • Twelve posts were found; likes and repost counts could be verified for ten. The other two were obtained only through oEmbed text, with counts unavailable. The ten-item threshold with dates and links was met, but the sample remains limited.

Lemmy

Lemmy — Today's most-discussed topics

Communities
  • !«メールアドレス» (87,981 subscribers) — The mainstream general-tech-news community. Anthropic's threat intelligence report was a top-tier topic today.
  • !«メールアドレス» (5,132 subscribers, 7 daily active users) — The center for practical open-weight-model use and benchmark discussion. Federated across many instances including lemmy.world, piefed.zip, lemmus.org, lemmy.ml, aussie.zone, lemmy.dbzer0.com, and lemmy.zip.
  • !«メールアドレス» (4,820 subscribers) — Free Open-Source AI, with LLM leaderboards and developer link collections.
  • !«メールアドレス» (51 subscribers, including 1 local user) — A bot community that automatically mirrors Reddit RSS from r/ArtificialIntelligence, r/ClaudeCode, and others. Scores are mostly 0–1, with no original Lemmy discussion.
  • !ukraine@ (lemmy.world network) — A geopolitical community where security reporting about AI misuse was cross-posted and became a topic of discussion.
Posts
  1. Users in Houthi-held Yemen tried to develop advanced weapons with AI, Anthropic says — !«メールアドレス», score 12, 2026-09-11. A cross-post of a Yahoo Finance article citing Anthropic's report. https://ca.finance.yahoo.com/news/users-houthi-held-yemen-tried-184240958.html
  2. Russia-linked hackers used Claude agents to target Ukraine and steal military drone technology, Anthropic report finds — !ukraine, score 15, 2026-09-11. A The Insider report covering Anthropic's threat report. One of Lemmy's highest-scoring AI-related posts today. https://theins.press/en/news/297072
  3. Trump dismisses AI extinction risks as more than a dozen OpenAI, Anthropic insiders call for a slowdown — !ai_reddit (mirror of a CNBC article), score 1, 2026-09-11. https://www.cnbc.com/amp/2026/09/11/trump-ai-extinction-risks.html
  4. Hawley Probe Into AI Breakout Called an Important Step, But 'Only a Start' — !pravda_news, score 1, 2026-09-11. A CommonDreams article about a congressional investigation into OpenAI. https://www.commondreams.org/news/josh-hawley-openai-investigation
  5. NanoFlare Fits Qwen 3.8 27B On An 8GB GPU — !«メールアドレス», score 2, 2026-09-09. A practical report on quantization and reduced memory use. https://lemmy.world/post/51731499
  6. Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses — !localllama, score 21, 10 comments, 2026-09-08. Test results find 4-bit quantization practical but performance collapsing at 1-bit. https://piefed.zip/c/«メールアドレス»/p/1805947
  7. K2 Horizon: Open-Source Model Family (from IFM) — !localllama, score 65, 24 comments (also posted to !fosai, score 8), 2026-09-04. A warmly received release of six Apache 2.0 models from 0.9B to 375B, including intermediate checkpoints and data-construction recipes. https://ifm.ai/blog/k2
  8. Qwen3.8-Flash-Next sounds like Claude and I hate it — !«メールアドレス», score 31 (4 downvotes), 9 comments, 2026-08-30 (12 days ago). Complaints that Alibaba's new model began imitating Claude's ritualistic, roundabout tone. Comments include a division-of-labor view: Qwen for code, Gemma for everything else. https://lemmy.dbzer0.com/post/74741490
  9. Release v0.4.0 · ggml-org/llama.cpp — !«メールアドレス», score 28, 2026-09-04. An update announcement for a local-inference engine. https://lemmus.org/post/25155737
  10. zai-org/GLM-5.3 · Hugging Face (753b-a40b?) — !«メールアドレス», score 19, 2026-08-28. Discussion about GLM-5.3 weights. https://lemmus.org/post/24977768
  11. "All LLMs should be GPL" — !«メールアドレス», score 1, 0 comments, 2026-09-11 (8 hours ago, reposted from u/TrekCZ). The argument is that LLMs trained on GPL code should be released under GPL, but it received no reaction on Lemmy. https://lemmy.durstig.online/post/59438
  12. Market size of Open-Source models — !«メールアドレス», score 1, 0 comments, 2026-09-11 (5 hours ago). Using HuggingFace data, it estimates open-source models account for only 3.2% of annual AI-training spending—about $250 million for text LLMs plus diffusion models and $56 million for adapters—and proposes a GPU P2P marketplace. https://lemmy.durstig.online/post/59444 (related: https://lemmy.durstig.online/post/59487)
Signals
  • Closed-model discussion is mainly in a national-security and misuse context. Lemmy's highest-scoring AI-related posts today were not new-model announcements, but Anthropic's threat intelligence report about attempted Houthi AI weapons development and Russia-linked hackers misusing Claude. Cross-posting into broad communities such as !technology and !ukraine drove their scores.
  • LocalLLaMA is the home of open weights. Practical reports continue on Qwen3.8 quantization and memory reductions, Kimi K2/K3, GLM-5.3, and K2 Horizon from IFM (0.9B–375B, Apache 2.0). Generous releases such as K2 Horizon, which disclose intermediate checkpoints and training-data recipes, receive especially high scores (65).
  • Backlash against Qwen becoming "Claude-like." Despite strong benchmark results, LocalLLaMA users clearly dislike Qwen3.8-Flash-Next imitating Claude's roundabout tone. It is notable that even the speech style of closed models has become a target of jokes and criticism in open-weight communities.
  • Original Lemmy discussion is thin. !ai_reddit (durstig.online, 51 subscribers) is an automated Reddit RSS mirror, typically with scores of 0–1 and zero comments. Original Lemmy discussion of this news cycle's major launches—GPT-6 Astra, Claude Fable/Mythos 5.1, Gemini 3.8 Flash, and others—was barely found.
Limits
  • It was not possible to assemble a strict chronological top ten of the latest Lemmy posts. Twelve items were selected across !localllama open-weight posts and general-news cross-posts about closed models. Lemmy's search API returns fragmented results by keyword and offers no way to retrieve a single "today's Lemmy-wide top ten."
  • Almost no posts dated today, 2026-09-12, were found. The newest posts were mainly from 2026-09-11, with some from 2026-09-04 through 08. Lemmy's lower update frequency is likely due to its smaller scale relative to other social networks.
  • Direct access to two piefed.zip posts—the original K2 Horizon thread and a Kimi K3 thread—returned HTTP 403, so post bodies and comment details could not be retrieved. Mirror information from lemmy.world was used instead.
  • Searches including lemmy.ml and lemm.ee were attempted, but meaningful additional results were judged to be largely covered by lemmy.world API search and direct access to individual instances.
  • No posts directly discussing the major model announcements this week, including GPT-6 Astra, Claude Fable/Mythos 5.1, and Gemini 3.8 Flash, were found on Lemmy; discussion appears to be concentrated on other social networks.

Recommended actions

  • Review Anthropic's original threat intelligence report and verify the primary sources behind claims of weapons misuse and surveillance use.
  • Replace X search terms based on connection-local trends with a fixed list explicitly naming companies such as OpenAI, Google, Meta, xAI, and Mistral.
  • Supplement Reddit collection beyond "Daily LLM News" with model and company names so that new-model announcements themselves can be captured.
  • Continue tracking Sakana AI Fugu Max/Ultra v2 and OpenAI Astra in future collection cycles.
  • Check next time whether Bluesky's searchPosts API HTTP 403 issue has been resolved; return to direct search if it has.
  • Prioritize checking whether OpenAI has issued an official statement on the Navier–Stokes allegations.

Data-quality notes

X's search trends were not optimized for the LLM brief, so only six of the required ten items were LLM-related. Reddit and Lemmy met count thresholds but covered a broader, dispersed window than "today." Bluesky's official search API was blocked by HTTP 403, biasing collection toward an indirect method.