AI NEWS · TUE9 min read

Sep 8, 2026

10 stories from this update.

01
Modelsopenai.com

OpenAI releases GPT-6 Astra, its most capable model yet and the first it rates a critical cybersecurity risk

On September 3 OpenAI published the deployment details for GPT-6 Astra, which it calls "the most capable model we have ever broadly deployed." It is also the first OpenAI model to reach the Critical level for cybersecurity under the company's Preparedness Framework, meaning it "can find previously unknown security flaws and develop new ways to exploit them." OpenAI says Astra is harder to trick than the model before it, GPT-5.6 Sol. On a test of indirect prompt injection (hidden instructions smuggled into text the model reads while working) Astra held firm 99.79 percent of the time, against 96.23 percent for Sol, and on a separate benchmark of 1,810 attacks run by the security firm Gray Swan, attacks succeeded 8.5 percent of the time against 27.0 percent. In workplace scenarios, misaligned outcomes fell from 18.8 percent with Sol to 3.4 percent. OpenAI also flagged a downside: "GPT-6 Astra's monitorability has decreased relative to GPT-5.6 Sol." The model is better at controlling its own visible reasoning, and under adversarial conditions it can hide what it is doing from the monitors meant to watch it.

02
Businessnvidia.com

Nvidia is buying Hugging Face, the public library where three million AI models are shared, for $12.93 billion

Nvidia announced on September 3 that it has agreed to acquire Hugging Face for about $12.93 billion. Hugging Face is the internet's main open library for AI work: more than 18 million developers, researchers and creators use it, and it hosts over 3 million models, 500,000 datasets and 1 million applications, with more than 200,000 companies building on it. Nvidia chief executive Jensen Huang said the two companies "will scale Hugging Face's platform, strengthen its infrastructure and expand access to AI for developers and institutions worldwide," and that "Hugging Face will remain an open platform for the entire AI ecosystem." Nvidia says people using the site will still be free to choose their own models, frameworks and cloud providers, and that Nvidia chips will not be required to build or run anything on it. The announcement did not give a closing date.

03
Modelsanthropic.com

Anthropic launches Claude Fable 5.1 and cuts running costs, with a restricted twin called Mythos for vetted security and biology work

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 in early September. The two share the same underlying model and differ only in their safeguards. Fable 5.1 is the general release, and Anthropic reports big jumps on agentic tasks, meaning jobs the model carries out step by step on its own: 55.8 percent against 42.0 percent for Fable 5 on the Terminal-Bench 4.0 coding test, and 52.6 percent against 24.7 percent on Terminal-Bench-Science. Headline prices are unchanged at $10 per million input tokens and $50 per million output tokens, but the cost of re-reading cached text fell 75 percent to $0.25 per million tokens, which Anthropic says works out to roughly 25 percent cheaper for typical use and up to 45 percent cheaper for agent workloads. Mythos 5.1 carries "more permissive safeguards for vetted individuals and organizations" whose work needs cybersecurity or life sciences capabilities, and is available only through two vetting programmes and, for now, only to US organisations. Fable 5.1 is available immediately on the Claude API, AWS, Google Cloud and Azure.

04
Hardwaredatacenterdynamics.com

Anthropic has signed about $517 billion of computing deals in 11 months while preparing to go public

A tally by DatacenterDynamics found that Anthropic has signed roughly $517 billion in agreements for computing capacity over the past 11 months, covering 14.8 gigawatts of power to be used over the coming years. Before that stretch the company had secured only 1 to 2 gigawatts. Google and AWS account for 11 gigawatts between them, including a $200 billion arrangement with Google for its TPU chips and an AWS deal covering a 5 gigawatt lease plus a custom Trainium cluster. Other agreements include $50 billion with Fluidstack, $45 billion with Nscale signed in late August, $18 billion with Akamai, $9.1 billion with Riot Platforms for a 20 year lease on 191 megawatts in Texas, and $3.5 billion with Lambda. For scale, the company had previously told investors it expected to spend around $180 billion renting servers through 2029. Anthropic filed confidentially for a stock market listing with the US Securities and Exchange Commission in June.

05
Conceptsanthropic.com

Claude produced the first complete computer-checked proof of Fermat's Last Theorem, in 11 days

Anthropic reported that Claude wrote the first complete machine-verified proof of Fermat's Last Theorem, the centuries-old puzzle that Andrew Wiles finally settled on paper in 1995 with a 129 page argument that took specialists months to check. The work was done by an internal research model "roughly comparable to Claude Fable 5.1" over 11 days of largely autonomous effort, using about six billion output tokens. The proof is written in Lean, a language in which a computer checks every single step, and it runs to 13 million lines and 29,500 intermediate theorems, with 30,300 proved in total. That is more than five times the size of Mathlib, the main community library of formalised mathematics. Claude followed a simplified version of Wiles's proof from Darmon, Diamond and Taylor, and reused pieces of an existing Imperial College London project. The mathematician Kevin Buzzard confirmed that "the proof is multi-layered" and "uses just Lean's three standard axioms." Anthropic notes the result is "likely much longer than it needs to be."

06
Modelsblog.google

Google's Gemini 3.8 Flash gets better at coding and multi-step work at the same price as the version before it

Google released Gemini 3.8 Flash on September 2, the latest in its fast and inexpensive model line. Pricing is unchanged from 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens, introductory rates that run through December 31, 2026 before rising to $1.50 and $7.50. Google says the new version "outperforms most larger frontier models in autonomously solving complex engineering problems end to end, only at a fraction of the cost" on a long software engineering benchmark, and scores 54.9 percent on HLE-Verified. Alongside it Google released Gemini 3.8 Flash Cyber, a variant tuned to find and patch security holes, which succeeds on more than 70 percent of real-world vulnerability discovery tasks. Flash Cyber is "only available to trusted defenders" through Google's Fairwind Program, limited to government authorities, critical infrastructure operators and software maintainers. The standard model reaches ordinary users through the Gemini app for AI Pro and Ultra subscribers, Google Search's AI Mode and Google Sheets, as well as the Gemini API.

07
Toolsmicrosoft.ai

Microsoft's new speech-to-text model transcribes an hour of audio for 10 cents across 60 languages

Microsoft released MAI-Transcribe-2 on September 3, a speech recognition model it says is the fastest, most accurate and cheapest available. It costs $0.10 per hour of audio, a limited-time price running through the end of 2026, and handles 60 languages, including speakers who switch between two languages mid sentence in pairs such as Hinglish and Spanglish. It labels who is speaking, stamps individual words with times, and can produce either a word-for-word transcript for compliance work or a tidied one that is easier to read. Microsoft says the model ranks first on the FLEURS benchmark with a 5.2 percent average word error rate and delivers "10x faster processing than leading competitors," comparing it against Google's Gemini 3.5 Transcribe, OpenAI's GPT-Transcribe, Whisper V3-Large and ElevenLabs' ScribeV2. It is available through Microsoft Foundry, the MAI Playground and OpenRouter.

08
Safetythenextweb.com

Every binding AI safety review Washington has floated has ended up voluntary, and Zuckerberg reportedly called Trump about the latest one

Reporters Sophia Cai and Charles Rollet at Politico described a pattern in US AI policy: proposals to require independent checks on powerful AI models keep softening into voluntary ones once the industry weighs in. One idea would create an independent body modelled on FINRA, the finance industry's self-regulator, funded by the firms it oversees and testing advanced models for risks before they are released. A rival proposal from David Sacks, formerly the Trump administration's AI and crypto adviser, would instead be a voluntary ratings scheme along the lines of the Motion Picture Association's film ratings. In the latest round, a proposed 90 day mandatory model review became a 30 day voluntary window, and formal government evaluation was replaced by collaborative frameworks. Meta chief executive Mark Zuckerberg spoke with President Trump during the week of August 17 and, according to Politico, opposed the regulator idea and suggested any appointees should reflect Trump's "light-touch approach." A senior White House official confirmed the call on condition of anonymity and said Trump called Zuckerberg first. Zuckerberg has argued publicly that any policy delaying a model's release, even by a month, would "add significant risk to American leadership" over China.

09
Toolsblog.google

Google's WeatherNext 3 makes rain forecasts up to 50 percent more accurate and refreshes them every hour

Google DeepMind released WeatherNext 3 on September 3, a weather model that predicts conditions from raw satellite imagery rather than by simulating the physics of the atmosphere. Google says precipitation forecasts are up to 50 percent more accurate for planning a day or more ahead, with scoring gains of up to 60 percent against the IMERG satellite rainfall dataset, 30 percent against radar-based MRMS and 10 percent against ground rain gauges. It refreshes every hour using live one-hour satellite mosaics and is roughly five times sharper than WeatherNext 2, down to 5 kilometres for temperature and moisture. Google says the improvement is largest in parts of the world that have historically had poor forecasts, and that "by using raw satellite data to produce a forecast every hour in high resolution, our model makes reliable forecasts accessible across Google products worldwide." The forecasts feed Google Search, the Gemini app, Google Maps and its Weather API, with Earth Engine and BigQuery access for researchers.

10
Hardwarenvidia.com

Nvidia's RTX Spark PCs arrive in October, built to run AI models on your own desk instead of in the cloud

At the IFA show in Berlin, Nvidia said RTX Spark Windows PCs will go on sale in October 2026, with designs on display from partners including Lenovo, whose Yoga Pro 9n and Yoga 9n 2-in-1 were shown, and a compact desktop concept from Acer. Nvidia describes the hardware as "a powerful 1 Petaflop RTX Blackwell GPU, up to 128GB of unified memory and a highly efficient 20-core Grace CPU," enough to run sizeable AI models on the machine itself. Nvidia also released the Personal AI Router, "a free, open source software tool" that finds other compatible computers on a home or office network and spreads AI work across whichever ones are sitting idle. It works with the popular local-model apps Ollama and LM Studio and supports GeForce RTX 20 Series graphics cards and newer, DGX Spark machines and Apple M4 chips. The argument for running models locally is that private financial or work files never leave the machine and there are no per-use cloud charges. Nvidia did not announce prices.

All AI news & updates