KD Agentic · AI Daily Digest, October 10, 2026. Seven stories today: Microsoft shipping a dedicated decision model, Anthropic turning Claude into a BI and motion tool, JetBrains upgrading its open coding model, Kodiak putting autonomous trucks on the busiest US border crossing, Arena nearly doubling its valuation in ten months, TypeSafe closing one of the largest Series A rounds of the year, and Google stripping free Gemini users down to a single model.
1. Microsoft ships Decision-1, a model that only makes choices
Microsoft CEO Satya Nadella announced Microsoft-Decision-1 on Friday, a new model that does one thing: given a set of fixed options, it picks one and attaches a calibrated probability to the choice. It is post-trained on Alibaba's open source Qwen3.5-9B, runs in a single pass, and supports yes/no questions, multiple choice, and scoring against a rubric, including rating AI answers or agent actions. The model is live in Microsoft Foundry's model catalog now, with OpenRouter access planned, and Microsoft says future versions will be retrained on its own MAI models and OpenAI models.
The benchmark claims are aggressive, and Microsoft is careful to label them as its own tests. Across 36 benchmarks and nearly 150,000 questions it reports the highest accuracy in the comparison set, running 4.5 times faster than Quyet-1.0-Large and roughly 35 times faster than GPT-6 Sol. Stability numbers matter as much as speed here: across eight kinds of perturbations, including rephrasing the same request and shuffling option order, verdicts changed only 1.3 percent of the time on average, and shuffling options produced zero flips. Pricing is set at $0.042 per million input tokens with output free, which puts it in the same band as the decision-model tier that formed over the past month.
The internal deployment examples show where Microsoft thinks the value is. The Xbox research team used Decision-1 to classify more than 10,000 player feedback items, matching GPT-6 Sol quality while running over 14 times faster and about 200 times cheaper, and the Copilot team uses it to score response quality at roughly 100 times the speed of a comparable model. The pitch is that agent workflows chain dozens of small judgments, and at 100 milliseconds of added latency per step a 20-step flow burns two seconds of pure waiting. Routing, intent detection, moderation, labeling, and choosing an agent's next action are the named targets, and the category now has Microsoft, TypeSafe's Jev, Liquid AI's d1, and OpenAI's Decisions API in public beta.
— Microsoft · Satya Nadella · Seeking Alpha
2. Anthropic turns Claude into a BI dashboard and an animation studio
Anthropic shipped two new features on Thursday, both in beta. Claude Dashboards connects to a company's data platforms, including BigQuery, Snowflake, Databricks, Amazon Redshift, ClickHouse, and Salesforce, and builds monitoring boards that update automatically from natural language questions. Every figure can be clicked open to inspect the underlying query, Claude can explain how a number was calculated, and each chart shows when the data last refreshed. Boards can be pushed into Amplitude, Grafana, Hex, Mixpanel, and PostHog for deeper analysis, with Looker and Tableau support coming, and the feature is open to all paid plans.
Claude Motion takes a different approach to video-adjacent content: instead of calling a video generation model, Claude writes code that animates your existing text, charts, shapes, and images into a short clip. A quarterly report becomes a 30-second explainer for an all-hands, and because every element is code-driven, each word, number, and duration stays editable before exporting to MP4. Type /motion in a conversation to use it. The feature is limited to Team and Enterprise plans for now, and the practical distinction from Sora-style generators is that Motion presents content you already have rather than fabricating footage.
Alongside the launches, Docs, Slides, and Design graduated from beta to general availability for every plan including free users, after users created more than 45 million documents, slides, and designs during the test period. The graduation batch adds simultaneous editing between teammates and Claude, sharing outside the organization with admin controls, PowerPoint and PDF exports that preserve layout, slide conversion into editable Google Slides, and mobile editing, plus CMEK key management for enterprises. One cleanup note for teams: the standalone Claude Design site at claude.ai/design closes on December 14, and chat history and project comments will not migrate over.
— Anthropic · IT之家
3. JetBrains upgrades Mellum with reinforcement learning for agentic coding
JetBrains released Mellum 2.1 on Friday, the second iteration of its open coding model family, keeping the 12B mixture-of-experts architecture with 2.5B active parameters and the Apache 2.0 license. The headline change is in the training pipeline: reinforcement learning moved from a short finishing stage to the main body of post-training, with new data mixed in across mathematics, competitive programming, science, tool use, and software engineering.
The infrastructure behind the RL stage is the more interesting detail. JetBrains built an internal reinforcement learning environment platform that spun up millions of sandboxes spanning thousands of environments during training, and filtered open datasets beforehand to drop broken tests, unverifiable answers, and badly calibrated task difficulty. The resulting model can explore a codebase, edit files, and check its own changes, and JetBrains says the largest gains show up in agentic coding, where Mellum 2.1 can identify the root cause of a failing test, draft a fix, and verify the result.
Performance stays close to Mellum 2 on architecture, so speed is unchanged, with multi-token prediction cutting response latency by about 1.6x for single requests. Under identical evaluation settings, throughput at high load approaches twice that of Qwen3.5-9B, with gains across coding, competitive programming, math, tool calling, and general knowledge. The model is on Hugging Face for local or self-hosted deployment, and GGUF builds for llama.cpp, Ollama, and LM Studio plus a vLLM multi-token-prediction speculative decoding component are planned next.
— JetBrains · IT时代网
4. Kodiak puts autonomous trucks on the Dallas-Laredo corridor
Kodiak AI and cross-border carrier Charger USA announced autonomous freight service on the 435-mile lane between Charger's terminals in Dallas and Laredo, Texas, per an official release dated October 8. Trucks equipped with the Kodiak Driver haul refrigerated and dry freight for consumer packaged goods and food and beverage customers, the first delivery ran on September 8, and a safety driver is behind the wheel until Kodiak completes its safety case for the lane, after which the companies plan driverless operation. It is Kodiak's first route serving Laredo, the highest-volume commercial land port of entry in the United States.
The lane choice is the story. The Port of Laredo handled more than $350 billion in international trade in 2025, close to 40 percent of all US trade with Mexico, and consumer goods freight on this corridor moves in steady retail replenishment flows where predictable transit times and longer operating hours are worth real money. Charger's president Andy Khera framed the combination of its in-house transportation management system with the Kodiak Driver as a nearly fully automated ecosystem that gives customers capacity continuity while getting drivers home every night.
Kodiak runs a Driver-as-a-Service model that deploys its system on customer-owned trucks, and its driverless commercial service in the Permian Basin has operated with nobody in the cab since December 2024, so the Laredo lane extends a proven configuration rather than debuting one. The company's long-haul Autonomy Readiness Measure, which tracks how much of its safety case is materially complete, reached 96 percent at the end of September, and Kodiak plans to launch driverless long-haul service by the end of 2026. Charger also signed an autonomous trucking agreement with Aurora for the same lane in late July, which makes the corridor an early proving ground where two driverless stacks will run in parallel.
— Kodiak AI · FreightWaves
🔗 Kodiak IR Press Release · FreightWaves
5. Arena nearly doubles its valuation to $3.1 billion in ten months
Arena, the crowdsourced AI leaderboard that grew out of UC Berkeley's LMArena project, announced on Thursday a $200 million Series B at a $3.1 billion valuation, co-led by Lightspeed Venture Partners and Khosla Ventures with Salesforce Ventures, 01 Advisors, Dell Technologies Capital, Endeavor Catalyst, and existing investors a16z and Felicis participating. The round lands ten months after a $150 million Series A at $1.7 billion, and the re-rating tracks revenue: annualized revenue tripled from $30 million in January to $100 million by June, driven by AI Evaluations, the paid analytics product it launched for labs and enterprises.
The timing reflects a shift in how the industry thinks about measurement. Labs have caught models gaming static benchmarks, and enterprises want help picking models for their own workloads rather than trusting leaderboard rank. Arena's own framing is blunt: AI is advancing faster than our ability to evaluate it, and static tests break down once models recognize they are being tested. The platform reports 350 million total sessions, 62 million votes across text, vision, code, search, video, and image, and 7 million sessions in Agent Arena, its agentic evaluation product, in under five months.
The round shipped with a new product line that puts Arena in the safety business: the Alignment Index, which scores models on unauthorized actions, false attribution, and deceptive completion, the failure modes that matter once agents act on people's behalf. On its preliminary numbers OpenAI models take the top five places, with Claude Opus 5.5 and Claude Fable in sixth and ninth. There is an obvious tension in a company funded by the ecosystem it grades, and Arena is now selling judgment as infrastructure to the same labs it measures, which is either a neutrality problem or the whole business model depending on who you ask.
— Arena · TechCrunch
🔗 TechCrunch · Arena
6. TypeSafe raises $870 million for the judgment-model layer
TypeSafe AI, the startup behind the Jev decision model, confirmed on Friday that it raised $870 million at a $7.5 billion valuation in a Series A led by Andreessen Horowitz, with Sequoia Capital, existing investor DCVC, and angel investors participating, and a16z's Martin Casado joining the board. The company published the news in its own characteristically blunt blog post, noting that it finds fundraising announcements boring and listing what the money buys instead: more machine-native models, more infrastructure, and the enterprise features customers have been asking for.
The numbers around Jev are the reason a Series A reached this size. The model, launched September 15, returns typed probabilistic judgments instead of generated prose, with responses in 70 to 500 milliseconds, input priced at $0.042 per million tokens and output free. TypeSafe claims roughly a third of the Fortune 500 are already using Jev, that the model passed one million users within days, and that it has saved customers millions of dollars in production, though founder Diogo Almeida's own X post put the figure at 29.4 percent and the company declined to name customers. The claims are the companies' own and remain lightly verified.
The funding tells you where a slice of AI capital is going: away from another conversational model and toward a decision layer that sits beneath agents, routing systems, moderation pipelines, and other production software. Almeida previously worked at OpenAI on the methods behind ChatGPT, and TypeSafe argues that abandoning text generation gives it speed and cost advantages LLMs cannot match on classification work. The bet has a clear failure mode, since calibrated confidence has to translate into measurable accuracy and safe escalation paths in real deployments, and the valuation gives the company little room to be merely fine at that.
— TypeSafe · Bloomberg
🔗 TypeSafe Blog · Bloomberg via TSN
7. Google cuts free Gemini down to Flash-Lite
Google's updated support documentation confirms that starting October 9, personal accounts without a Google AI subscription can only use Gemini Flash-Lite, with the faster Flash and the more capable Pro models removed from the free tier. AI Plus subscribers keep Flash-Lite and Flash but lose Pro, while AI Pro and AI Ultra retain all three, and AI Pro gains Deep Think, the maximum parallel reasoning mode previously reserved for higher plans. Free-tier changes took effect on the day, with AI Plus transitions rolling out per account by email notice.
Google frames the change as model access moving behind subscription tiers, a different mechanism from the compute-based usage limits it introduced in May, which cap how much you can use but not which model you can reach. A usage limit is temporary; a model-access restriction means a free user's debugging question that Pro would have handled now goes to Flash-Lite or nowhere. For casual use, summarizing, rewriting, and routine questions, little changes, but the ceiling for demanding work is now set by the plan you pay for rather than the quota you have left.
The move is a clean test of consumer AI pricing power. Google is betting that enough free users hit the capability wall on coding, complex reasoning, and long documents to convert, while AI Pro gains a headline feature to justify its own step up. It also tightens the ladder across the industry's consumer offerings in the same week that OpenAI, Anthropic, and Microsoft all pushed product changes, and it makes the free tier of a flagship assistant noticeably thinner, which competitors will not miss.
— Google · ezone.hk
KD Agentic · AI Daily Digest

Top comments (0)