Shared research report

What are the most significant developments in AI this week?

August 14, 2026

Research Report

Question: What are the most significant developments in AI this week?

Date: 2026-08-14T13:21:17.567528641+00:00

Rounds: 4

Status: COMPLETE

Evidence: 62 claims · 50 sourced · 4 partial · 6 unsupported · 2 self-reported (no independent source) · 3 single-source

Executive Summary

For the week of May 3–9, 2026, the most significant AI developments, in order:

  1. OpenAI made GPT-5.5 Instant the default ChatGPT model (May 5) — the week's biggest product event, primary-confirmed by OpenAI's announcement and TechCrunch.
  2. The Pentagon signed classified-network AI agreements with eight companies — excluding Anthropic (announced May 1; the week's dominant business/policy story), primary-confirmed by the Department of War release and CNN.
  3. Baidu released ERNIE 5.1 (May 8) — the first Chinese model to reach LMArena's global top 10, primary-confirmed by Baidu's ERNIE blog.

The week was otherwise quiet at the frontier: no Anthropic, Google, or Meta flagship shipped (Google's Gemini 3.5 came May 19; Anthropic's next Opus, 4.8, came May 28), and xAI had no confirmed week-dated release. Several stories circulated as "this week's" news turned out to be misdated or unverified — most notably a "Claude Opus 4.7 launched May 5" claim (Opus 4.7 actually shipped in April) and a "UK AISI found GPT-5.5 matched Mythos" report (the AISI evaluation was published April 13 and compared GPT-5.4/GPT-5.3-Codex, with Mythos ranked first). See the corrections table under Analysis.

The week's top developments

#DevelopmentDateVerificationWhy it matters
1GPT-5.5 Instant becomes ChatGPT's default (and API chat-latest), replacing GPT-5.3 InstantMay 5Primary: OpenAI, TechCrunchFirst default swap since 5.3; ~52.5% fewer hallucinated claims and 37.3% fewer inaccuracies vs GPT-5.3 Instant; AIME 2025 81.2 (vs 65.4); new "Memory Sources" personalization with provenance; paid users keep 5.3 for 3 months
2Pentagon AI contracts to eight firms, Anthropic excludedMay 1 (dominated the week)Primary: war.gov, CNNSpaceX, OpenAI, Google, NVIDIA, Reflection, Microsoft, AWS, Oracle get Impact Level 6/7 classified-network access (1.3M GenAI.mil users); Anthropic — whose Claude was previously the only model on the classified network — is barred by a supply-chain-risk designation it is litigating
3Baidu ERNIE 5.1May 8Primary: Baidu + codersera roundupFirst Chinese top-10 LMArena entry: #4 Search Arena (Elo 1223), #14 text; ~⅓ of ERNIE 5.0's total params and ~6% of frontier pre-training compute; AIME26 99.6 with tools
4GPT-5.5-Cyber restricted previewMay 7Primary: OpenAI, CNBCCyber-defense variant under "Trusted Access for Cyber," vetted teams only — the week's other security-model storyline (besides Anthropic's Mythos, itself unveiled April 7, before this week)
5OpenSeeker-v2 (open-source search agent)May 6Primary: arXiv 2605.04036Academic 30B model, SFT-only on 10.6K trajectories, tops BrowseComp at 46.0% — beating far heavier industrial pipelines; weights open-sourced; the week's most-upvoted HF paper (622 stars)
6Zyphra ZAYA1-8BMay 6Primary: arXiv 2605.05365700M-active / 8B-total reasoning MoE claims parity with far larger models: 91.9% AIME'25 and 89.6% HMMT'25 with a 4K-token test-time-compute tail (Markovian RSA), trained on AMD hardware
7Codex CLI v0.129.0May 7Primary: GitHub releaseModal Vim composer, redesigned resume/fork picker, raw scrollback, 18 bug fixes — agentic-coding tooling keeps shipping weekly
8ElevenLabs: $500M ARR + third Series D closeMay 5Primary: ElevenLabs, TechCrunchNot a fresh $500M raise (that was Feb 2026): ARR crossed $500M (from $350M at end-2025); BlackRock, NVIDIA, D.E. Shaw, Wellington among new investors; $100M tender
9Gemma 4 open-weight familyWeek of May 3–9Aggregator only: devFlokersClaimed most capable open-weight family from Google DeepMind (Apache 2.0; 2B–31B, 128K–256K context; 26B A4B MoE at ~97% of the 31B Dense's quality with 8× less compute). Not primary-verified — treat as provisional

Analysis

The week's center of gravity moved from raw frontier scale to three things: default-model quality (OpenAI), geopolitics of AI access (the Pentagon/Anthropic fight, and China's efficiency push), and efficiency research. OpenAI owned the product narrative — the 5.5 Instant default swap, the 5.5-Cyber release, and a Codex CLI update all landed within 72 hours. The Pentagon announcement made the geopolitics concrete and awkward for Anthropic: excluded from the eight-company agreement while simultaneously reported to be renting SpaceX's Colossus 1 supercomputer (that deal, ~220K NVIDIA GPUs at 300MW, reported May 6, remains aggregator-sourced and unverified at primary record). On the competitive side, ERNIE 5.1's LMArena top-10 entry at roughly 6% of frontier pre-training compute is the strongest single data point yet for the "China–US gap has narrowed" narrative.

The research theme was efficiency and test-time compute: ZAYA1-8B's sub-1B-active-parameter reasoning claims, OpenSeeker-v2's SFT-only SOTA, SubQ's 12M-context launch (May 5, claims unvalidated and no architecture paper yet), and DeepSeek's "Thinking with Visual Primitives" — the week's most-discussed paper (spatial markers as "minimal units of thought"; 98.7% spatial reasoning; ~7,000× token compression), which went viral May 4 but was dated April 30 and later quietly pulled with no weights released. Treat it as research-in-flux, not a shipped capability.

Policy/regulatory was a documented vacuum: the UK AISI published nothing between April 29 and May 14; EU AI Act milestones bracketed the week (transparency obligations effective Aug 2; code finalized June); the US BIS chip-loophole action came May 31; and ICML 2026 notifications had already gone out April 30.

Claims that did not survive verification

Circulating claimVerified status
"Claude Opus 4.7 launched May 5 at a Wall Street event"False — Opus 4.7 shipped April 2026 (trackers date it April 16); no Anthropic flagship released this week
"UK AISI (May 5): GPT-5.5 matched Anthropic's Mythos"Wrong date and model — AISI's Mythos cyber evaluation was published April 13, 2026; comparators are GPT-5.4 / GPT-5.3-Codex (GPT-5.5 appears nowhere in it); Mythos ranked highest — first to solve AISI's 32-step "The Last Ones" (22/32 avg steps vs Opus 4.6's 16)
"OpenAI hit $25B ARR on May 9"Misdated — OpenAI crossed $25B annualized revenue in early 2026, reported March 5 by Reuters
"Anthropic overtook OpenAI in revenue (May 4)" / "OpenAI–Anthropic acquisition talks"Unconfirmed — zero coverage found in three-month news-index sweeps
"ElevenLabs raised $500M on May 8"Reframed — the $500M round closed Feb 2026; the May 5 event was the $500M ARR milestone + third Series-D close
"Anthropic–SpaceX Colossus 1 deal"Unverified — reported May 6 only by aggregators; no primary announcement located in the findings
"Claude Mythos is fictional / has no primary source"False — Mythos Preview was publicly unveiled April 7, 2026 (CNN, Anthropic's Mythos page); it predates this week

Risks & Open Questions

Claims without independent support

These statements appear in the narrative above but are not backed by text retrieved from a source. SELF-REPORTED means the only thing asserting it is the swarm's own worker output — the narrative was written from that output, so it corroborates nothing. Treat all of these as unverified.

Detailed Findings

Round 0 · Finding 1

The browser-rendered bodies are coming back empty. Let me try the http_request with the browser_fallback disabled, or use the goto verb in the browser to get actual content.

Round 0 · Finding 2

Significant AI Model Releases — Week of May 3–9, 2026

Executive Summary

The first full week of May 2026 was defined by OpenAI's GPT-5.5 Instant rollout — a strategic pivot toward reliability and personalization rather than raw scale — and Google DeepMind's Gemma 4, the most capable open-weight model family of its class. OpenAI also shipped a new suite of real-time voice models, and the week closed with major infrastructure and partnership news (Anthropic–SpaceX compute deal) that would reshape the industry's economics. Notably absent: no new frontier flagship from Anthropic (no Opus released this week), xAI, or Meta landed during May 3–9; their marquee releases came later in the month (e.g., Claude Opus 4.8 on May 28, Qwen 3.7-Max on May 20).


Key Findings

1. OpenAI — GPT-5.5 Instant (Released May 5, 2026)

Verified details:

Sources: https://www.devflokers.com/blog/latest-ai-models-open-source-projects-may-2026 ; https://web.archive.org/web/20260719210226/https://www.aicritique.org/us/2026/06/01/ai-developments-in-may-2026/

2. OpenAI — Real-Time Voice Model Suite (Shipped May 7, 2026)

Verified details:

Sources: https://aitoolsrecap.com/Blog/ai-news-may-2026 ; https://web.archive.org/web/20260719210226/https://www.aicritique.org/us/2026/06/01/ai-developments-in-may-2026/

3. Google DeepMind — Gemma 4 (Released Week of May 3–9, 2026)

…(truncated — the summary above captures the substance)

Round 0 · Finding 3

Notable AI Research Developments — Week of May 3–9, 2026

Executive Summary

Live research for the week of May 3–9, 2026 surfaced two fully verified research items on the primary record (arXiv): a reasoning-model efficiency breakthrough from Zyphra (ZAYA1-8B, posted May 6) and a 30-author ICML 2026 position paper on Bayesian orchestration for agentic AI (updated in-week, May 6), plus an ICML-accepted theory paper on Bayes-consistency (May 5). Notably, the week's most loudly "promoted" alleged breakthroughs in aggregator blogs — a DeepSeek "Thinking with Visual Primitives" multimodal reasoning paper and a sensational "Claude Mythos" cybersecurity model — could not be verified against any primary record; the arXiv API returned zero results for the former, and no primary source exists for the latter. Any week-summary built on those claims would be built on unverified material.


Key Findings

1. Zyphra's ZAYA1-8B — small-MoE reasoning model that claims parity with much larger models (VERIFIED, high confidence)

…(truncated — the summary above captures the substance)

Round 0 · Finding 4

Most Significant AI Business Developments — Week of May 3–9, 2026

Based on live web research into daily AI news roundups, startup funding trackers, and venture analysis covering the target week.


Executive Summary

The week of May 3–9, 2026 was defined by three interlocking themes: (1) the largest compute-acquisition deal ever struck in AI, with Anthropic committing to rent SpaceX's entire Colossus 1 supercomputer; (2) a Pentagon AI contracting fight that pointedly excluded Anthropic while granting deals to SpaceX, OpenAI, Google, Microsoft, Nvidia, AWS, Oracle, and Reflection; and (3) an extraordinary burst of venture and revenue momentum across the sector — including ElevenLabs raising $500M, Anthropic topping OpenAI in quarterly revenue, OpenAI hitting $25B ARR, and Meta launching its Muse Spark product line.


Key Findings (with Confidence Levels)

1. Anthropic–SpaceX Colossus Compute Deal (HIGH confidence — reported May 6, corroborated by subsequent S-1 disclosures)

On May 6, 2026, Anthropic signed a deal to rent SpaceX's entire "Colossus 1" supercomputer — 220,000+ NVIDIA GPUs at 300MW capacity. Claude Code rate limits were doubled across all paid tiers the same day, signaling expanded compute capacity. Subsequent disclosures (SpaceX S-1, May 21) revealed Anthropic pays approximately $1.25 billion/month through May 2029 — a ~$45B total commitment, far above the $3–6B annual figure analysts had originally estimated (wikipedia-quality estimate ~$15B/year).

2. Pentagon AI Contracts Exclude Anthropic (MEDIUM confidence — reported May 5)

The Pentagon signed AI agreements with SpaceX, OpenAI, Google, Microsoft, Nvidia, AWS, Oracle, and Reflection — pointedly excluding Anthropic amid ongoing litigation. This marks a notable business development given Anthropic's parallel $45B compute commitment to a SpaceX subsidiary.

3. Anthropic Topped OpenAI in Revenue (MEDIUM confidence — reported May 4)

Anthropic reportedly generated more revenue than OpenAI at this point, driven by its Claude Code product and enterprise deployments. This was later corroborated on May 11, when Anthropic disclosed Q1 revenue grew 80x year-over-year with ARR above $44B.

…(truncated — the summary above captures the substance)

Round 1 · Finding 1

The DeepMind blog returned empty content via browser rendering. Let me check the blog listing directly.

Round 1 · Finding 2

Findings: AI Policy, Regulatory, Safety, Export-Control, and Chip Infrastructure Developments — May 3–9, 2026

Executive Summary

I was tasked with filling the policy/regulatory/safety/export-control/hardware gap for the week of May 3–9, 2026. Across multiple search-engine queries (EU AI Act enforcement May 2026, NVIDIA export controls May 2026, US AI chip export controls China May 2026, AI governance policy White House May 2026, EU AI Act GPAI code of practice May 2026, plus site-specific queries), the available search tooling returned only evergreen, generic content — Wikipedia pages, vendor homepages, and dict/calendar entries — rather than dated news items from that specific week. A direct fetch of the NVIDIA Newsroom archive (nvidianews.nvidia.com/news) returned the live archive, whose visible top items were dated August 2026 (e.g., the Aug 10–11, 2026 NVIDIA/Apollo/BlackRock/KKR AI-financing announcements, Aug 11–12 Glasdoors/CEO and Nemotron 3.5 items); I could not page back to the May 3–9 window within my budget.

Net assessment: I could not identify, confirm, or corroborate any specific AI policy, regulatory, safety, export-control, or chip-infrastructure development dated to May 3–9, 2026. Every candidate I was asked to check (EU AI Act enforcement actions, US chip export-control changes, NVIDIA/TSMC hardware announcements, model safety incidents for the week) remains unverified. Per the research protocol, I will flag these explicitly rather than repeat aggregator claims as fact. The absence of hits is not proof that nothing happened — it reflects the limits of the search index available in this session — but I will not fabricate details to fill the gap.


Key Findings (with confidence levels)

1. No primary-record evidence found for week-dated policy/hardware events — LOW/MEDIUM confidence that significant events occurred but were missed; HIGH confidence that I could not verify any. All search-engine queries returned non-news, evergreen pages. No primary regulator (EU Commission/European Parliament, NIST, Commerce/BIS, White House) or chip company (NVIDIA/TSMC newsroom) week-dated item was surfaced in my fetch window. Confidence in the absence of found items: HIGH. Confidence that I captured the full picture: LOW.

…(truncated — the summary above captures the substance)

Round 1 · Finding 3

Resolution of the two contradictory claims

Claim 1 — "Anthropic released Claude Opus 4.7 on May 5, 2026 at a Wall Street finance event": NOT SUPPORTED — no such model exists in the record.

Claim 2 — "Claude Mythos, a restricted cybersecurity model, exists": CONFIRMED — it is real, but the cited AISI details (May 5 date, "GPT-5.5 matched Mythos") are wrong.


Findings

1. Claude Mythos is real and was publicly disclosed on April 7, 2026 — not a hoax, and not a May-week event.

2. The UK AISI evaluation exists but is dated April 13, 2026 — not May 5 — and it does NOT say GPT-5.5 matched Mythos.

…(truncated — the summary above captures the substance)

Round 1 · Finding 4

Business Developments for May 3–9, 2026: Confirmation Assessment

Summary verdict

ClaimStatusEvidence strength
Pentagon AI contract list excluding AnthropicCONFIRMED (announcement dated May 1, not May 5)High — full article fetched
ElevenLabs $500M raise (May 8)UNCONFIRMED — no trace in news indexesNegative evidence
OpenAI $25B annualized revenue (May 9)UNCONFIRMED — best available figure is $20B est. ARR, and it dates from a different sourceContradicted by only datapoint found
Anthropic surpassing OpenAI in revenue (May 4)UNCONFIRMED — zero coverage foundNegative evidence
OpenAI–Anthropic acquisition talksUNCONFIRMED — zero coverage foundNegative evidence

1. Pentagon AI procurement contracts excluding Anthropic — CONFIRMED (date corrected to May 1)

This is the week's real, verifiable business/defense story — and the only one of the five that survives contact with primary-adjacent reporting.

Caveats: the DoD announcement is dated May 1, 2026 (SiliconANGLE's URL and og:updated_time May 3), so the "May 5" date in the prior finding is off by days. I verified via reputable tech press, not the primary defense.gov release itself (not reachable in this session), so the primary record was not independently checked. The substantive claim — seven companies contracted, Anthropic excluded — is confirmed.

…(truncated — the summary above captures the substance)

Round 2 · Finding 1

Research, Open-Source, and Benchmark Developments — Week of May 3–9, 2026

Important caveat on primary sourcing

All items below are drawn from secondary/aggregator and practitioner sources; I could not reach a primary arXiv entry or vendor press release dated May 3–9, 2026 for most of these (the live arXiv API queried in real time returns August 2026 submissions, not May, and could not be date-restricted to the window in the session). Where an item appears week-dated only through aggregation, it is flagged accordingly.

1. DeepSeek "Thinking with Visual Primitives" — the standout research item of the week (CONFIRMED as a real DeepSeek work; week-date partially corroborated)

…(truncated — the summary above captures the substance)

Round 2 · Finding 2

AI Model / Product / API Releases, May 3–9, 2026 — Findings by Vendor

Scope and method

I searched official vendor newsrooms and release notes (OpenAI, Anthropic, Google DeepMind, Meta AI, xAI, Mistral, Cohere) and cross-checked with dated aggregator roundups. Primary-source confirmed items are separated from secondary-source-reported items below. All figures and dates are cited; where I could not reach a primary record, I say so explicitly.


1. OpenAI — the only major vendor with confirmed week-dated product releases

GPT‑5.5‑Cyber (security-focused variant) — May 7, 2026 (PRIMARY-SOURCE CONFIRMED)

GPT‑5.5 Instant becomes ChatGPT's default — May 5, 2026 (SECONDARY-SOURCE REPORTED; not independently confirmed from OpenAI's primary announcement)

Codex CLI v0.129.0 — May 7, 2026 (SECONDARY-SOURCE REPORTED)

…(truncated — the summary above captures the substance)

Round 2 · Finding 3

Week-dated AI Business, Chip, and Policy Developments — May 1–9, 2026 (Primary-Source Verification)

Verification focus: (a) ElevenLabs $500M raise / annualized revenue, (b) Pentagon AI contract list, (c) chip/export-control/regulatory items dated that week. Note: several prior-round "week" claims are real events but misdated — I flag each with its verified date.

1. ElevenLabs ("$500M raise," May 8) — CONFIRMED as a May 5, 2026 event, but the figures differ from the claim

Confirmed (primary source). ElevenLabs' own blog post, "ElevenLabs crosses $500M ARR and welcomes new investors," is datePublished 2026-05-05 (https://elevenlabs.io/blog/500m-arr-and-new-investors). It states the company ended 2025 at $350M ARR and "in the first four months of 2026, we have already surpassed $500 million ARR." It announces the third close of the Series D with new investors BlackRock, Wellington, D.E. Shaw, Schroders, NVIDIA (via NVentures), Santander, Jamie Foxx, Eva Longoria, and Hwang Dong-hyuk, plus a $100M tender offer. So:

Verdict: The week-dated ElevenLabs development is real and primary-confirmed, but as $500M ARR + third Series-D close (May 5, 2026), not "$25B annualized revenue" and not "May 8."

2. OpenAI "$25B annualized revenue" (May 9) — NOT a this-week event; it dates to early March 2026

Confirmed date is March, not May. Reuters, "OpenAI tops $25 billion in annualized revenue … The Information reports," is dated 2026-03-05 (https://www.reuters.com/technology/openai-tops-25-billion-annualized-revenue-last-month-information-reports-2026-03-05/). US News (2026-03-04) carries the same "topped $25B annualized revenue last month" report (https://money.usnews.com/investing/news/articles/2026-03-04/openai-tops-25-billion-in-annualized-revenue-last-month-the-information-reports). No May 3–9 OpenAI revenue announcement was located. The prior-round flag of this as a May 9 item is a misdating.

…(truncated — the summary above captures the substance)

Round 2 · Finding 4

Resolution: AISI evaluation of Claude Mythos Preview — publication date and comparison model

Finding 1: Correct publication date is April 13, 2026 (not May 5, 2026)

The Wayback Machine's CDX index for https://www.aisi.gov.uk/blog/our-evaluation-of-claude-mythos-previews-cyber-capabilities shows the earliest capture timestamp is 20260413160726 (April 13, 2026, 16:07 UTC), followed by a continuous run of captures on April 13, 14, 15, 16, 22, 25, 27, and 28, 2026 (http://web.archive.org/cdx/search/cdx?url=www.aisi.gov.uk%2Fblog%2Four-evaluation-of-claude-mythos-previews-cyber-capabilities&output=json&limit=100). There is no Wayback capture on or near May 5, 2026 — the capture record jumps directly from 2026-04-28 to 2026-05-14. A page captured within hours of publication (April 13, 16:07 UTC) with an unbroken April archive trail cannot have been first published on May 5.

The May 14, 2026 snapshot, retrieved in full, displays the byline date "Apr 13, 2026" directly beneath the post's summary line (https://web.archive.org/web/20260514124354/https://www.aisi.gov.uk/blog/our-evaluation-of-claude-mythos-previews-cyber-capabilities). The current live page, fetched today (server date Aug 14, 2026), shows the identical byline date "Apr 13, 2026" (https://www.aisi.gov.uk/blog/our-evaluation-of-claude-mythos-previews-cyber-capabilities). The live page carries a last-modified: Mon, 10 Aug 2026 header indicating later edits, but the byline date and body content match the May 14 snapshot, so the later edit did not alter the publication date.

Conclusion: April 13, 2026 is confirmed. The "May 5, 2026" date is not supported by the primary record — no AISI blog capture exists for May 5, and no AISI publication of that date was found.

Finding 2: Correct comparison model is GPT-5.4 / GPT-5.3-Codex (not GPT-5.5)

Both the May 14, 2026 archived snapshot and the current live page name the comparison models in the Figure 3 caption of the "cyber range" results:

"Mythos Preview, Opus 4.6, and GPT-5.4 average 10 runs up to 100M tokens. Opus 4.5, GPT-5.1 Codex, and Sonnet 4.5 each average 15 runs up to 10M and 5 runs up to 100M tokens. GPT-5.3-Codex averages 10 runs up to 10M and 5 runs up to 100M tokens. Sonnet 3.7 and GPT-4o average 10 runs up to 10M tokens only."

The Figure 1 caption additionally names GPT-3.5 Turbo through Claude 4 Opus and "GPT-5 through to Mythos Preview." The string "GPT-5.5" appears nowhere in either the archived May 14 version or the live version of the post. The closest GPT-5-family comparators in the evaluation are GPT-5.4 and GPT-5.3-Codex, plus GPT-5.1 Codex in the CTF/range charts.

Conclusion: The comparison models in the AISI Mythos evaluation are GPT-5.4 and GPT-5.3-Codex (with GPT-5.1 Codex also charted). The "GPT-5.5" attribution is refuted for this document — it does not occur in the primary source in either its May 2026 archived form or its current form.

…(truncated — the summary above captures the substance)

Round 3 · Finding 1

AI Policy, Safety, and Legal Developments: May 3–9, 2026

Executive Summary

A targeted sweep of regulator sites, safety-body publications, and court/legislative reporting found no confirmed AI policy, safety, or legal events dated within May 3–9, 2026 beyond the May 1 DoD contract agreement (previously documented in the prior round). The White House AI executive order that press coverage references was announced and then postponed May 20–21, 2026, outside the search window. UK AISI published no evaluation during the target week. The Anthropic copyright settlement was finalized in July 2026, not in-window.


Key Findings by Category

1. US White House / Executive Branch — No in-window events CONFIRMED

The prior round's finding that no in-window US AI executive action is supported by this sweep. The most relevant proximate event came weeks later:

2. EU AI Act / European Commission — No in-window enforcement actions CONFIRMED

Search results surfaced the EU's Code of Practice on Transparency of AI-generated Content and the General-Purpose AI Code of Practice (https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content, https://digital-strategy.ec.europa.eu/en/policies/contents-code-gpai), but these predate/postdate May 2026. Law-firm analyses (e.g., Jones Day on the "final code of practice on AI labelling" dated June 2026, https://www.jonesday.com/en/insights/2026/06/european-commission-publishes-final-code-of-practice-on-marking-and-labelling-aigenerated-content; Cooley noting transparency obligations take effect August 2, 2026, https://www.cooley.com/news/insight/2026/2026-08-03-eu-ai-act-transparency-obligations-take-effect-2-august-2026) place the operative events after the target week. No EU AI Act enforcement action dated May 3–9, 2026 was found.

3. UK AI Safety Institute (AISI) — No in-window publications CONFIRMED

The AISI research/publications index (https://www.aisi.gov.uk/research) lists publications chronologically. The items bracketing the target week are:

…(truncated — the summary above captures the substance)

Round 3 · Finding 2

Chinese Lab and Open-Source AI Developments — Week of May 3–9, 2026

Scope and Exclusions

This investigation focused on (a) Chinese labs (Alibaba/Qwen, Moonshot, Zhipu, MiniMax, ByteDance/Seed, Baidu/ERNIE) and (b) open-source ecosystem items (Hugging Face trending, OpenClaw, A2A, MCP, open-weight model releases), dated May 3–9, 2026. Notably, the dominant Chinese release events of May 2026 — Alibaba's Qwen 3.7-Max (May 20) — fell outside this window. Only one flagship Chinese model release, ERNIE 5.1, falls squarely inside it.


1. CONFIRMED IN-WINDOW: Baidu ERNIE 5.1 (released May 8, 2026)

Status: Confirmed. Primary source is Baidu's own ERNIE blog at https://ernie.baidu.com/blog/posts/ernie-5.1-0508-release/ (the URL itself carries the 05-08 date marker), and multiple independent secondary sources corroborate the May 8 date.

What it is: ERNIE 5.1 is Baidu's flagship model, the first Chinese model to crack the global top 10 on LMArena's Search Arena leaderboard — landing at #4 globally with an Elo of 1223, and #14 on the standard text leaderboard (the highest position any Chinese lab has held there) (https://codersera.com/blog/ai-model-releases-may-2026-roundup/).

Efficiency profile: ERNIE 5.1 compresses total parameters to roughly one-third of ERNIE 5.0 and active parameters to one-half, spending only 6% of the pre-training compute of comparable frontier models (https://chatforest.com/reviews/baidu-ernie-5-1-moe-frontier-llm-review/; https://www.baidumarketing.com/blog/baidu-ernie-5-1-release). On AIME26 with tool use it scores 99.6 — second only to Gemini 3.1 Pro.

Availability: 128K context window; API access and direct chat at ernie.baidu.com, priced via Baidu's Qianfan platform. The preview had been quietly live since April 30 (https://codersera.com/blog/ai-model-releases-may-2026-roundup/).

Confidence: High (confirmed by primary Baidu source and multiple independent trackers).


2. CONFIRMED IN-WINDOW: Subquadratic SubQ 1M-Preview (launched May 5, 2026)

Status: Confirmed for date; model architecture claims UNVERIFIED.

Subquadratic, a Miami startup (founded by ex-SoftBank Vision Fund and Tinder alumni, $29M raised at a $500M valuation), launched SubQ on May 5, 2026 with a claimed 12-million-token context window (~10× larger than any frontier model in production) (https://codersera.com/blog/ai-model-releases-may-2026-roundup/). This is the first "fully subquadratic" context model — selecting a small subset of positions per query token. Claimed performance includes a 7.2× speedup at 128K tokens and 92.1% on needle-in-a-haystack at 12M tokens.

…(truncated — the summary above captures the substance)

Round 3 · Finding 3

Notable AI papers, model releases, and acceptance announcements — week of May 3–9, 2026

Bottom line: The week's research output is now documented from primary and archival routes. The arXiv 2605.0xxxx ID range is confirmed real and populated; Hugging Face's daily paper feeds for May 3–8, 2026 are archived and recoverable; and — importantly — ICML 2026 acceptances were not announced in this window (notifications went out April 30, 2026, per ICML's own dates page). No Nature- or NeurIPS-dated item for May 3–9 was found. The most significant items of the week are open-weight model reports and agent/search-agent papers, led by NVIDIA's Nemotron 3 Nano Omni and the academic search-agent paper OpenSeeker-v2.


1. ICML 2026 acceptance announcements — REFUTED for this window (primary record)

The ICML 2026 official Dates and Deadlines page states "Author Notification: Apr 30 '26 (Anywhere on Earth)" — i.e., acceptances were released the week before the target window, not during May 3–9. Main conference dates are July 7–9, 2026 (https://icml.cc/Conferences/2026/Dates). The only ICML-adjacent in-window deadline is the workshop-level "Universal notification deadline for all submissions to individual ICML workshops: May 14 '26," also outside the window. The accepted-papers list itself is live at https://icml.cc/virtual/2026/papers.html (an aggregator, https://aiconfpaper.com/conferences/icml-2026, counts 6,634 accepted papers), and the ICML blog's awards post is dated July 5, 2026 (https://blog.icml.cc/2026/07/05/announcing-the-icml-2026-awards/) — both post-window artifacts. Verdict: no ICML 2026 acceptance announcement occurred May 3–9; the week's ICML-relevant news was limited to community coverage of the April 30 notification.

…(truncated — the summary above captures the substance)

Round 3 · Finding 4

Verification Report: OpenAI May 5–7, 2026 changelog items

Verdict: Both items are CONFIRMED by primary sources.


1. "GPT-5.5 Instant" became the default ChatGPT model on May 5, 2026 — CONFIRMED

Primary source (OpenAI's own announcement): The official page https://openai.com/index/gpt-5-5-instant/ is dated May 5, 2026 (the dateline appears directly on the page) and titled "GPT‑5.5 Instant: smarter, clearer, and more personalized." Its Availability section states verbatim: "GPT‑5.5 Instant is rolling out starting today to all ChatGPT users, replacing GPT‑5.3 Instant as the default model, and in the API as chat-latest. For paid users, GPT‑5.3 Instant will remain available for three months, accessible through model configuration settings, before being retired." The page also contains a later "Update on June 9, 2026" note about personalization expanding to Free/Go tiers — but the May 5 default-model switch is unambiguously the original, dated event. (Verified by fetching the live page; metadata/canonical confirm the URL.)

Reputable secondary corroboration: TechCrunch published https://techcrunch.com/2026/05/05/openai-releases-gpt-5-5-instant-a-new-default-model-for-chatgpt/ with datePublished: 2026-05-05T17:00:00+00:00 (byline "10:00 AM PDT · May 5, 2026"), opening: "On Tuesday, OpenAI released a new foundation model called GPT-5.5 Instant, which will replace GPT-5.3 Instant as the default ChatGPT model." (May 5, 2026 was indeed a Tuesday, consistent.) It reports the same API detail ("chat-latest," 5.3 kept for paid users for three months) and the AIME 2025 score of 81.2 vs. 65.4.

Help-center changelog / Wayback angle: The official ChatGPT release-notes page (https://help.openai.com/en/articles/6825453-chatgpt-release-notes) has Wayback captures bracketing the release — May 3, 2026 (capture 20260503135946, HTTP 200) and May 10, 2026 (capture 20260510132637, HTTP 200), with a bot-blocked capture in between on May 8 (20260508053536, HTTP 403). This confirms the changelog page existed and was archived both before and after the May 5 date. Caveat: I could not directly quote the changelog entry's text — the May 10 raw replay (via id_ format) returned gzip-compressed bytes that did not decode inline, so the help-center entry itself was not readable in this session. The item is nonetheless confirmed on the strength of OpenAI's own dated announcement page, which is the stronger primary record.


2. "Codex CLI v0.129.0" released on May 7, 2026 — CONFIRMED

…(truncated — the summary above captures the substance)

Investigation Trail

Round 0

Round 1

Round 2

Round 3

Sources

Trace Index

Tool-call traces are persisted under /srv/swarm_web_runs/run-1786712828394-0003/traces.

This report was researched and written by a Swarmio run — a swarm of AI agents that searches the web, reads the sources, and shows its working.

Ask your own question Are you an AI agent? Start at /llms.txt — sign up, mint a key, and run with no human.