AI news on August 3 ranged from ambitious mathematics claims to a warning about fabricated vulnerability reports entering formal data systems. Here are eight developments, with the source boundary attached to each one.
1. OpenAI published ten mathematics and theoretical computer science advances
OpenAI says an internal version called Astra solved or materially advanced ten long-standing open problems. According to the company, the model generated mathematical arguments, humans helped organize them into papers, and the model then formalized the work into Lean certificates. OpenAI estimated that the tokens used to find the solutions would cost about $2,000 at Sol API prices.
This remains an OpenAI publication, not a settled academic verdict. Correctness, novelty, and the status of each result still require review by the relevant research communities.
Source: https://openai.com/index/ten-advances-in-mathematics/
2. Qwen3.8-Max was unveiled, but its open weights are not available yet
Alibaba announced Qwen3.8-Max with 2.4 trillion total parameters and 95 billion active parameters. The published specification includes text, image, and video input plus a one-million-token context window. The weights are expected next week.
Those specifications and performance claims are vendor statements. The release of the weights has not happened yet, and benchmark or long-running task results should not be described as independently verified.
Source: https://qwen.ai/blog?id=qwen3.8
3. A third-party test put DeepSeek V4-Flash at a very low token price
Reuters cited Artificial Analysis measurements that priced V4-Flash at about $0.14 per million input tokens and $0.28 per million output tokens. The reported average cost was about $0.03 per benchmark task.
This is a third-party benchmark and pricing snapshot. Token prices alone do not establish the total cost, reliability, or value of a model on a real workload.
Source: Reuters, August 3, 2026, reporting Artificial Analysis measurements. The source brief did not include the article URL.
4. JFrog says 54 reviewed advisories appeared to be LLM fabrications
JFrog investigated a new GitHub repository that published more than 50 vulnerability advisories. The researchers concluded that 54 advisories appeared to be LLM-generated fabrications. The reports cited missing functions, unrelated code lines, or proof-of-concept inputs that failed before reaching the alleged vulnerable code. Several reports still entered downstream systems and received high or critical severity metadata.
The finding applies to this batch. It does not show that the entire CVE system is untrustworthy. JFrog's strongest evidence came from checking source code and PoCs, not from relying on an AI detector.
Source: https://research.jfrog.com/post/sqlite-critical-cves-or-llm-slops/
5. Google released Gemini Robotics ER 2 for high-level robot reasoning
Gemini Robotics ER 2 is available through the Gemini API and Google AI Studio. Google says it supports continuous video progress tracking, self-correction, tool use, and coordination among multiple robots. Its role is high-level reasoning rather than direct motor control.
Google's enterprise platform remains in private preview. Public demonstrations and evaluations are company-led, and real performance depends on the robot hardware and control stack.
Source: https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-robotics-er-2/
6. The EU opened a call for up to seven AI gigafactories
The European Union launched a call for up to seven AI gigafactories. The plan aims to combine up to EUR 10 billion in public funding with at least EUR 20 billion in private investment.
This is a procurement and financing target. It is not completed computing capacity, and it does not mean EUR 30 billion has already been committed or received.
7. Microsoft reported more than 30 million paid Copilot seats
Microsoft's FY26 Q4 results said Microsoft 365 Copilot had more than 30 million paid seats. The same release said annual Azure revenue passed $100 billion for the first time.
Paid seats are not the same as monthly active users, retention, or standalone AI revenue. The figure is a sales measure, not a complete picture of usage depth.
Source: https://www.microsoft.com/en-us/Investor/earnings/FY-2026-Q4/press-release-webcast
8. Thinking Machines Lab released full Inkling-Small weights
Thinking Machines Lab released the full weights for Inkling-Small. The company describes it as a mixture-of-experts model with 276 billion total parameters and 12 billion active parameters, support for a one-million-token context, and native image and audio reasoning.
Full weights do not mean the model is easy to run on an ordinary computer. Vendor comparisons should also remain separate from independent evaluation.
Source: https://thinkingmachines.ai/news/inkling-small/
Together, these eight items show several different stages of the AI pipeline: research claims, model releases, prices, security data, robotics, infrastructure, enterprise sales, and open weights. The source boundaries matter because a company announcement, a third-party benchmark, an investigation, and a procurement target are different kinds of evidence.
Disclosure: AI tools assisted with drafting and source organization. The factual claims and limitations were checked against the linked source brief before publication.
Top comments (0)