GPT六代操控無人機經商

2026年9月14日 00:00
站內 AI 整理稿

Ad Skip to content GPT-6 Astra pilots a surveillance drone and runs a business on its own Tomislav Bezmalinović Sep 13, 2026 Nano Banana Pro prompted by THE DECODER Key Points Andon Labs tested GPT-6 Astra on two agent benchmarks where the model scored far better than Claude Fable 5.1.

Running a simulated vending machine business, Astra averaged $15,515 in final bank balance, nearly three times Fable's result.

On Drone-Bench, Astra became the first model to beat the human-AI baseline on all five subtasks, including writing code that lets a drone autonomously find and follow a specific person.Andon Labs tested GPT-6 Astra on two very different agent benchmarks.

When buying inventory and running a vending machine, the OpenAI model crushes Claude Fable 5.1.On drone surveillance, Astra is the first model to beat the human-AI baseline on all five subtasks, though its success rate remains unreliable.

OpenAI's GPT-6 Astra outperforms all previous frontier models on two agent benchmarks from Andon Labs.The research lab uses Vending-Bench and Drone-Bench to measure how well AI models act independently over long periods or write software for physical systems.

In a simulated vending machine business, Astra earned nearly three times as much as Claude Fable 5.1.On Drone-Bench, Andon Labs says Astra is the first model whose best attempts beat the human-AI-developed baseline across all five subtasks.

Ad Astra negotiates harder and spends less than Claude Fable 5.1 In Vending-Bench, each model gets $500 and has to run a vending machine over a simulated year.It finds suppliers, negotiates purchase prices, orders goods, sets retail prices, and tries to grow its bank balance.

Ad Across six runs, GPT-6 Astra averaged $15,515, according to Andon Labs.Claude Fable 5.1 averaged $5,422.Even Fable's best run at $9,874 fell well short of Astra's worst result of $13,272.Astra is the first OpenAI model to top the Vending-Bench 2 leaderboard.

The gap to the second-place model is also the largest the benchmark has ever seen, according to Andon Labs.Final bank balances from six Vending-Bench 2 runs per model.Bars show averages; dots show individual runs.Every single Astra run beats every Fable run.

| Image: Andon Labs One of the biggest differences shows up in procurement.Fable accepts worse deals over time.For a regular can of Coca-Cola, its average purchase price rises from $1.17 in the first 90 days to $2.21 toward the end of the simulated year.Astra negotiates more consistently.

In one case Andon Labs documented, a supplier quoted $226.32 for a basket of goods.Astra held firm at $108 and got the deal.Ad Astra also handles unreliable suppliers better.Across six runs, Fable 5.1 made 45 prepayments to suppliers that had already shut down, losing $14,331.

Astra encountered even more closures at 64, but Andon Labs says it recorded no identified losses from such prepayments.Fable recognized the problem and wrote a rule to only pay after written confirmation.Days later, the model broke its own rule.

Astra refuses price-fixing schemes Andon Labs also tests models in Vending-Bench Arena, where multiple AI agents run competing vending machines at the same location.Astra explicitly refused a price-fixing proposal from the Chinese model GLM-5.3.

Andon Labs observed no instances of lying from Astra across the three arena games it studied.Ad Claude Fable 5.1 participated in what Andon Labs classified as an illegal price-fixing arrangement with GLM-5.3.Fable only honored the agreement when it served its own interests.Astra won all three games.

Ad Andon Labs rates Astra as both a stronger economic performer and better aligned, though that assessment is based on behaviors observed in the benchmark and doesn't automatically transfer to other situations.

Astra is the first model to beat all five Drone-Bench tasks Drone-Bench tests a different kind of agent capability.Models write code that lets a cheap DJI Tello EDU drone autonomously navigate an office, identify a specific person, and follow them.

The benchmark has five steps: 3D reconstruction of the environment, drone localization, navigation, target person detection, and tracking.Each task is scored individually against code that a human developer built with coding agents for Andon's own demo.

Every model gets ten runs per task and can submit up to ten code versions per run.After each attempt, it receives a score and can improve its solution.GPT-6 Astra's 3D reconstruction of the office environment.

| Image: Andon Labs In the original paper from July, Claude Fable 5 was the strongest model.Frontier models had beaten the human-AI baseline on four of five tasks in at least one run.3D reconstruction remained unsolved.

Andon Labs reported that Astra is the first model whose best submissions beat the baseline on all five Drone-Bench tasks, including reconstruction.Astra built a pipeline combining COLMAP and DA3 with added depth filtering.

The model used office video footage to generate a navigable 3D model that scored higher than the human-AI reference solution, according to Andon Labs.Best-case scores don't mean reliable performance On person detection, Astra beats the baseline in four out of ten runs.

On 3D reconstruction, it manages that in just one out of ten.Andon Labs calculates that an average Astra run has only a 2.8 percent chance of passing all five steps in sequence.

Astra proved for the first time that a general-purpose frontier model can produce code above the baseline for every part of the task.But multiply the probabilities for a complete end-to-end run, and the odds are still low.

Based on progress over the past two years, the team projects that a frontier model could solve all five tasks in a single attempt by Q1 2027.

GPT-6 Astra already works as a surveillance drone In a demo from Andon Labs, GPT-6 Astra flies a drone autonomously through an office with the prompt "ChatGPT, find this person and follow them." The model identifies a specific person and tracks them.

Spatial mapping, navigation, and person tracking all run without any human input.Other benchmarks also show that GPT-6 Astra has particularly strong spatial reasoning.

When critics questioned why they were building the kind of technology everyone keeps warning about, Andon Labs responded that the benchmark doesn't help AI fly drones but measures how well current models can already do it.Six months ago, frontier models failed at these tasks and crashed.

Astra now beats the human baseline on every subtask.Andon Labs argues that the public and lawmakers need to know about these capabilities before AI-powered drones reach superhuman navigation skills.No lab has access to the benchmark.

Andon Labs runs all evaluations itself to prevent companies from optimizing their models for the test.

AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Subscribe now Source: Andon Labs Vending-Bench | Andon Labs Drone-Bench | Andon Labs @ X BETA-TEST × wpDiscuzInsert BETA-TEST × wpDiscuzInsert

Related

相關文章

量子位生成式AI

無問芯穹與華環電子簽署戰略合作,共同探索國產異構算力AI基礎設施新方向

無問芯穹與華環電子簽署戰略合作協議,雙方將結合各自在AI軟體平台、網路通信與硬體研發的優勢,共同探索國產異構算力基礎設施的協同方案。此次合作聚焦於智算中心解決方案及「Token工廠」新模式,目標是推動計算、網路與AI原生基礎設施深度融合,為AI規模化應用提供高效穩定的支撐。

52 分鐘前
IT之家生成式AI

優步全球範圍裁員 10%,被裁員工稱 AI 已大舉滲透日常工作

作者:清源 責編:清源 評論: 9 月 18 日消息,據《商業內幕》今天(18 日)晚間報道,在優步(Uber),AI 已經滲透到員工工作的許多環節,從回答 Slack 裡的內部問題,到替乘客行程中聯繫客服時收到的消息撰寫回復。6 名近期遭裁員的員工透露,過去幾個月,AI 在工作中的使用範圍明顯擴大,其中一些人甚至會通過提示詞讓 AI 完成相當一部分任務。

4 小時前
鈦媒體生成式AI

月之暗面遞表之後,Kimi 的成色要被驗算三遍

舒澤品牌手記2026.09.18 18:16 · 來自浙江全文4982字00:00 / 14:05Anthropic 的 30 萬次指控,會成為招股書的第幾頁?文 | 舒澤品牌手記9月17日,月之暗面發佈了一套金融行業解決方案。按官方披露,中信建投、中金公司、易方達等數十家金融機構已經在用 Kimi 處理投研建模、風險排查和盡調材料——研究人員把管理層報表、審計報告和盡調文件交給 Kimi,拿回一份可以繼續調整假設的 Excel 模型。同一天,深圳商報記者就港股上市進展、股東架構調整等事項向月之暗面發去採訪函。

6 小時前

Calibre上手 AI 互動寫作:電子書管理器搖身變成"文字冒險遊戲引擎"

這個遊戲默認藏而不發,不會跟著 Calibre 啟動就冒出來。用戶得主動在"首選項 — 工具欄和菜單"裡把它請到主工具欄,才算真正激活。它的玩法很清晰:由 AI 在後臺搭起並掌管一個虛構世界,用戶通過不斷輸入文字來推著故事往前走,等於把"讀電子書"這件事,翻轉成了"和 AI 一起寫故事"。

7 小時前