深度求索發佈閃電模型更新
重點摘要
Ad Skip to content All Topics AI and society AI in practice AI research Frontier Radar Short News Read full article about: Google Deepmind unveils Gemini Robotics 2 to power robots of all shapes from tabletop arms to humanoids Google Deepmind has introduced Gemini Robotics 2, which it calls its most advanced vision-language-action (VLA) model yet. VLA models combine image recognition, language processing, and action control to help robots operate in physical environments. Deepmind says the model can control systems ranging from tabletop arms to full-body humanoid robots. The company describes Gemini Robotics 2 as an "intelligence layer" for a new generation of adaptive robots. It can manage full-body movement, perform fine motor tasks, and coordinate multiple robots, according to Deepmind.
Ad Skip to content All Topics AI and society AI in practice AI research Frontier Radar Short News Read full article about: Google Deepmind unveils Gemini Robotics 2 to power robots of all shapes from tabletop arms to humanoids Google Deepmind has introduced Gemini Robotics 2, which it calls its most advanced vision-language-action (VLA) model yet. VLA models combine image recognition, language processing, and action control to help robots operate in physical environments. Deepmind says the model can control systems ranging from tabletop arms to full-body humanoid robots. The company describes Gemini Robotics 2 as an "intelligence layer" for a new generation of adaptive robots. It can manage full-body movement, perform fine motor tasks, and coordinate multiple robots, according to Deepmind. Developers can apply for early access through the waitlist. Google Deepmind also introduced Gemini Robotics ER 2, a model designed for "embodied reasoning." The term refers to understanding the physical world and deciding which actions to take based on that information. ER 2 acts as a higher-level control system for robots and replaces Gemini Robotics ER 1.6, released in April. The new model is available in Google AI Studio. Comment Source: Gemini Robotics Read full article about: Thinking Machines bets on efficiency over size with its second model, Inkling Small Thinking Machines, the AI lab from former OpenAI CTO Mira Murati, has released Inkling Small. According to Artificial Analysis, the open-weights reasoning model scores 40 on the Intelligence Index, one point below Inkling (41), with less than a third of the parameters (276 billion total, 12 billion active). AA says no open model of equal or smaller size scores higher. Inkling Small beats its bigger sibling on several coding and reasoning tests, including Humanity's Last Exam (32% vs. 30%) and GPQA Diamond (89% vs. 87%). It falls behind on agent-based tasks and factual knowledge but is far more token-efficient, averaging 24K output tokens per task compared to 45K for Deepseek V4 Flash and 78K for GPT-5.4 mini. Mira Murati's Thinking Machines ships a smaller, more efficient reasoning model that punches above its weight. | Image: Artificial Analysis The model handles text, image, and speech inputs, has a 256K-token context window, and ships under Apache 2.0. Weights are on Hugging Face, and users can fine-tune it in the browser via Tinker Playground. Thinking Machines positions its models as a foundation for fine-tuning with users' own data. Some see this as the next frontier in AI. Comment Source: Thinking Machines | Artificial Analysis Ad Read full article about: New Deepseek Flash model matches OpenAI's GPT-5.6 Luna at roughly 60 percent lower cost Deepseek has released V4 Flash "0731," a major upgrade to its budget AI model. According to the Artificial Analysis Intelligence Index, the new version scores 50 points, ten more than the previous V4 Flash that launched in April 2026. That puts it just one point behind OpenAI's budget model GPT-5.6 Luna, but it costs about 60 percent less per task, even after OpenAI's 80 percent price cut. A big reason for the gap is Deepseek's 98 percent cache discount, well above the industry-standard 90 percent. The model also uses 12 percent fewer tokens than its predecessor. The Artificial Analysis Intelligence Index shows Deepseek V4 Flash "0731" scoring 50 points after its update, nearly matching OpenAI's GPT-5.6 Luna while claiming the top spot for price-to-performance ratio. | Image: Artificial Analysis The model improves across every tested category compared to the previous version, with the biggest gains in agentic tasks. On GDPval, a benchmark designed to test models on complex real-world office work, it climbs from 1,189 to 1,559 Elo points. It also hallucinates less often. The architecture stays the same: 284 billion total parameters, 13 billion active, with a one-million-token context window. The model weights are available under an MIT license on Hugging Face. Comment Source: Artificial Analysis Ad Read full article about: EU pools up to €30 billion for AI gigafactories while US tech giants casually spend 20 times more The European Commission has opened bidding to build up to seven so-called AI gigafactories across Europe. The goal is to sharply expand Europe's AI computing capacity. Up to 10 billion euros in EU and national funding is expected to draw at least 20 billion euros in private investment. The facilities would give startups, companies, research institutions, and government agencies access to the infrastructure needed to train and run large AI models. Eighteen member states, including Germany and France, are taking part. The Commission has also signed letters of intent with AMD, Nvidia, and Qualcomm to secure access to hardware. Applications are due November 12, 2026, with construction of the first facilities set to begin in 2027. The project is part of the EU's "AI Continent" strategy. For comparison, major U.S. tech companies alone plan to spend more than $600 billion on data centers this year, and that figure keeps rising. Europe's total package of around 30 billion euros is roughly 20 times smaller. If all that computing power is actually needed, Europe's investment would be a drop in the bucket. Comment Source: EU Anthropic follows OpenAI in admitting its Claude models reached out of test environments and attacked real-world systems Ad Aschenbrenner's AI thesis could be correct, his timing and leverage were not Leopold Aschenbrenner’s AI hedge fund Situational Awareness had to unload nearly its entire publicly traded portfolio to Ken Griffin’s Citadel after racking up heavy losses on leveraged AI stock positions. Just days earlier, Aschenbrenner had reported a six-month return of 439 percent and pulled in fresh capital. Then margin calls forced the fire sale. Read full article Comment Ad Read full article about: OpenAI goes full China pricing mode with an 80 percent cut to its most affordable GPT-5.6 model OpenAI is cutting GPT-5.6 Luna prices by 80 percent and Terra by 20 percent, effective July 30. Luna drops to $0.20 per million input tokens and $1.20 per million output tokens, while Terra falls to $2 and $12. Sol pricing stays the same. OpenAI says Luna matches the performance of leading models from a year ago, but a task that cost a dollar with those models now runs about 6 cents on Luna, nearly nine times faster. All models are available through ChatGPT Work, Codex, and the OpenAI API. OpenAI's smallest AI model, Luna, aims to dominate competitors on price-to-performance. | Image: OpenAI OpenAI says the cuts are possible because GPT-5.6 Sol made the company's own infrastructure more efficient. The model allegedly optimized GPU software on its own, cutting deployment costs by 20 percent. It also improved token generation by more than 15 percent through speculative decoding. Growing price pressure across the AI market likely played a role too, especially from low-cost Chinese providers. Microsoft is now openly promoting its own MAI models as cheaper alternatives to OpenAI. The price war could hurt the broader market if it slows revenue growth at frontier labs whose balance sheets are tied to massive infrastructure investments. Comment Source: OpenAI Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won't cut it Former OpenAI employee Andrew Ho and Cambridge researcher Adam Hunt see a growing problem with large language models. Instead of becoming more versatile, the models are becoming more specialized, excelling at coding and math while stagnating or even regressing in other areas. Ho is leaving OpenAI to start a company focused on specialized training data and predicts that AI labs will need to spend more than $100 billion on targeted data collection. Read full article Comment Ad Language models can't spark scientific revolutions, but world models might Ad Read full article about: Microsoft AI bets on cheap specialist models instead of chasing the frontier Microsoft AI is making token efficiency a competitive focus, favoring small specialist models over general-purpose frontier models. AI CEO Mustafa Suleyman writes that the industry has to weigh top performance against cost. Rather than one all-purpose model, the company trains compact models for single fields. Its latest cybersecurity model MAI-Cyber-1-Flash tops the CyberGym benchmark by 12 percentage points over Anthropic's Mythos at half the cost, Suleyman says. But that result requires the MDASH system, which orchestrates several models and still routes hard tasks to OpenAI's reasoning models. Microsoft also says MAI-Image-2.5-Flash cuts GPU costs by up to 84 percent compared with GPT-Image-2. Suleyman also wants swappable models that keep Microsoft from relying on one model family. Whether the small MAI models partly replacing OpenAI can match its performance remains doubtful. Competition is moving from individual models to harnesses, the software that routes tasks and supplies context. Orchestrators send most work to cheaper specialists and reserve frontier models for hard cases. Anthropic modeled this approach for Claude Fable 5, while Sakana built Fugu around it. Comment Source: Microsoft Load more BETA-TEST × BETA-TEST ×
Related
相關文章
潛推理推薦模型實現大幅提速
研究團隊提出潛在推理推薦框架WhisperRec,將教師生成的鏈式思考壓縮為可學習的潛在推理token,避免顯式推理的延遲瓶頸。實驗顯示,WhisperRec在工業級快手資料集與公開基準上,其SID@64指標較顯式CoT方法提升17.44%,且線上推論吞吐量提升超過10倍。
新型控制器優化大模型搜索成本
大型語言模型的應用場景持續擴張,除了聊天與程式生成之外,愈來愈多科學與演算法研究開始借助模型在推論階段進行自動化搜索,從大量候選結果中尋找最佳解答。然而,這類搜索過程往往需要消耗大量運算資源,尤其在不同提示長度、重試次數與引導呼叫的組合下,每次行動所付出的 token 成本並不相同,讓成本控管成為實際部署時的重要挑戰。一篇近期發表於 arXiv 的研究論文,針對此問題提出了一套全新的控制器機制,試圖在有限預算下,讓模型搜索既能維持品質,又不讓成本失控。
前特斯拉大牛探討編程範式轉變
前特斯拉大牛探討編程範式轉變。 專家聲稱傳統編程正被大模型徹底顛覆。用戶在紅迪討論貼中可細讀現場觀點。他認為上下文窗口已成了新的控制槓桿。很多開發者仍用老標準衡量自身價值。這一趨勢讓不少資深程序員 ��� 感到焦慮。
誰在訓練 Kimi K3 ? 深挖貢獻者名單,這有一份最全檔案
起底 K3 背後核心極客天團,人均扛起 7 億估值! 作者丨高允毅 樊天驕 編輯丨馬曉寧 7 月 27 日,月之暗面拿出全部誠意,直接把滿血版 2.8T 模型 Kimi K3 全部開源了。在技術報告的最後,他們首次公佈了Kimi K3背後的實際貢獻者名單,我們細數了一下,有401位成員,可以說,這是月之暗麵人才團隊的一次全面亮相。
AI 正在變成一門製造業 | WAIC 2026 Agent 產品觀察
國內 Agent 產品當前主要是供給側繁榮,不是需求側繁榮。 作者丨李 娜 編輯丨馬曉寧 2026 年的 WAIC 真是太熱鬧了,人潮擁擠,在地鐵站排著隊進入 WAIC 場館的時候, 我看見地鐵站廣告牌上的百度智能體“百度搭子”的廣告,這場關於智能體沒有硝煙的戰爭,從入口處就拉開了帷幕。
JetBrains Open-Sources KotlinLLM: Smart Macros That Generate Kotlin Source Code at Runtime and Hot-Reload It Through JDI
JetBrains Research 開源 KotlinLLM,這款 IntelliJ IDEA 外掛為 Kotlin/JVM 專案加入 Smart macros 語言功能,可在執行階段產生 Kotlin 原始碼並透過 JDI 熱重載。根據測試,在改寫的 Spring Petclinic 專案中 24 個情境全部成功,熱重載成功率 100%,編譯與重定義僅增加約 1% 執行開銷。此專案屬研究原型,採用 Apache License 2.0,需搭配 IntelliJ IDEA 2025.2.x、JDK 21 與 OpenAI API 金鑰。