代理式編碼要取代初階工程師,必須滿足哪些條件?

2026年8月26日 14:20
站內 AI 整理稿

I read every major model release.Most of them ship a coding number.The number goes up.The conclusion everyone draws is that junior engineers are finished.I think that conclusion is being reached the wrong way.

People are reasoning from a benchmark score to a labor market outcome, skipping every step in between.So let me do it differently.Instead of asking “will agents replace juniors,” I want to ask what would have to be true for that to happen.

Then check each condition against the best evidence available.There are four.Three of them are not met.The fourth is the one that should worry you, because it does not require the other three.#mtp-fc-embed hr, #mtp-fc-embed p:empty, #mtp-fc-embed del, #mtp-fc-embed s { display: none !

important; } #mtp-fc-embed { background: transparent !important; border: 0 !important; margin: 28px 0 !important; padding: 0 !important; line-height: 1px !important; } #mtp-fc-embed iframe { display: block !important; width: 100% !important; border: 0 !important; background: transparent !

important; margin: 0 !important; padding: 0 !important; } @media (max-width: 640px) { #mtp-fc-embed { margin: 20px 0 !important; } } (function () { var f = document.getElementById('mtpFcFrame'); if (!f) return; window.addEventListener('message', function (e) { if (e && e.data && typeof e.data.

mtpFcHeight === 'number') { f.style.height = e.data.mtpFcHeight + 'px'; f.setAttribute('height', e.data.mtpFcHeight); } }, false); })(); Condition 1: Agents have to be reliable at the length of task a junior actually gets The best measurement we have here is METR’s time-horizon work.

They time human experts on real software tasks, then find the task length at which a model succeeds 50% of the time.The main result is that this horizon doubled roughly every seven months from 2019 to 2025.METR’s updated Time Horizon 1.

1 expanded the task suite by 34% and doubled the count of tasks running eight hours or longer.Independent readings of the 2024 to 2026 window suggest the doubling has since accelerated.The live leaderboard now puts frontier horizons in the hours.That sounds decisive.

Read the methodology and it stops being decisive.Two things are important: First, 50% is not a bar you can staff against.Kwa et al.also report an 80% horizon, and at any given moment it is dramatically shorter than the 50% figure.

In their data, frontier systems are near-perfect on tasks a human finishes in under four minutes and succeed less than 10% of the time on tasks that take a human more than four hours.

Second, and this is the part almost nobody quotes: METR says its tasks are deliberately self-contained and well-specified.

Their own framing is that a two-hour task should be read as what someone with no prior context could do in two hours, not what an experienced engineer familiar with the codebase could do.That is precisely the wrong shape.A junior engineer’s first six months are almost entirely context acquisition.

Which service owns this.Why that abstraction exists.Who to ask.The benchmark measures the one part of the job that has been stripped of the thing that makes it hard.

Condition 2: The benchmark has to measure the job In February 2026, OpenAI stopped reporting SWE-bench Verified and recommended others do the same.Their reasoning is worth reading in full, but two findings stand out.They audited a 27.6% subset of the dataset and found that at least 59.

4% of the audited problems had flawed test cases that reject functionally correct solutions.And they found contamination: frontier models could reproduce exact gold patches and verbatim problem details, indicating training exposure.State of the art had moved from 74.9% to 80.9% over six months.

The question OpenAI asked was whether the remaining failures reflected model limits or dataset properties.The answer was mostly dataset properties.Move to a harder, less contaminated set and scores fall off a cliff.

SWE-bench Pro was built for exactly this, and frontier performance on it sits far below the Verified figures the launch posts advertise.Newer suites like Terminal-Bench and long-horizon evolution benchmarks are being built for the same reason.I want to be careful here.

This is not “benchmarks are useless.” It is narrower and more damaging: the specific number that has been used for two years to argue juniors are obsolete was retired by the lab that created it, for reasons that make the number look better than reality.

Condition 3: The cost of verifying agent output has to fall below the cost of delegating to a person This is the condition I think gets ignored most, and it is the one with the cleanest experimental evidence.

METR ran a randomized controlled trial with 16 experienced open-source developers across 246 real tasks in their own repositories.AI allowed or disallowed at random.Screen recordings.Real work.Developers forecast a 24% speedup.Afterwards they estimated they had been 20% faster.They were 19% slower.

Two caveats, because I would rather you trust the rest of this piece.The tools were early-2025.The sample is small and specific: experienced developers on mature codebases they know well.This is not a universal productivity estimate and METR does not claim it is.

But the perception gap is the durable finding.People were wrong about the direction of their own productivity, under measurement.The wider data points the same way.

Stack Overflow’s 2025 survey of more than 49,000 developers found 84% using or planning to use AI tools, while 46% actively distrust the accuracy of the output against 33% who trust it.Only 3% report high trust.Among experienced developers, high distrust runs at 20%.

Google’s DORA research surveyed around 5,000 professionals and found 90% using AI at work and over 80% believing it lifted their productivity, while 30% report little or no trust in AI-generated code.DORA’s throughput finding improved from the prior year.Delivery instability did not.

Their conclusion is that AI is an amplifier: it magnifies what the organization already is.Put it together.Generation got cheap.Verification did not.Review capacity is now the constraint, and review capacity is senior engineer time.

Condition 4: Firms have to be willing to break their own senior pipeline Here is the uncomfortable part.Conditions 1 through 3 describe whether the substitution works.Condition 4 describes whether firms will attempt it anyway.They are not the same question, and the second one is already answered.

Stanford’s Digital Economy Lab tracks ADP payroll data covering roughly one in six American workers.Their Canaries work finds that employment for 22 to 25 year olds in the most AI-exposed occupations, software development among them, has diverged sharply from older workers in the same occupations.

The shortfall measured 15% at the July 2025 data vintage.As of June 2026 it is 19%.The live dashboard shows the adjustment running through reduced hiring rather than separations.Nobody is being fired.The door is closing.

The mechanism the revised paper proposes is the most interesting finding in any of this.Employment fell among young workers in occupations that lean on codified knowledge, the kind you can learn from documentation and standardized procedure.

It rose among experienced workers in occupations that lean on tacit knowledge, acquired through practice, mentorship and repeated exposure to real situations.Stanford is careful that these are descriptive patterns, not causal estimates.Take that seriously.

But if the mechanism holds, notice what it implies.Codified knowledge is what a junior arrives with.Tacit knowledge is what a junior is supposed to acquire, by doing the codified work under supervision until the tacit part sinks in.

We are automating the apprenticeship and keeping the requirement for what the apprenticeship produced.What I actually think Agentic coding is not replacing junior engineers.It is replacing the tasks we used to hand junior enginee

Related

相關文章

突發,Opus 5.1被曝本週發佈,最強AI智能體恐大洗牌

獨家獲悉,一款代號為 Opus 5.1 的 AI 智能體模型預計將於本週正式發布。消息一出,迅速在人工智慧業界引發劇烈討論,許多開發者與投資人紛紛預測,這款新模型可能顛覆現有智能體領域的競爭格局,導致「最強 AI 智能體」的稱號面臨重新洗牌。 根據多位知情人士透露,Opus 5.1 並非現有主流模型的常規升級版本,而是從底層架構到推理邏輯都進行了大幅改寫。雖然具體技術細節尚未公開,但業界普遍推測該模型在複雜任務分解、多步驟自主決策以及跨工具調用能力上,將達到前所未有的水準。

剛剛
IT之家AI Agent

2026 中國電信研究院 AITMark 獎項出爐,華為小藝獲 AI 手機智能體綜合評分第一

首頁 IT圈 最會買 設置 日夜間 隨系統 淺色 深色 主題色 黑色 投稿 訂閱 RSS訂閱 收藏 軟媒應用 App客戶端 要知App 軟媒魔方 業界 手機 電腦 測評 視頻 AI 蘋果 iPhone 鴻蒙 軟件 智車 數碼 學院 遊戲 直播 5G 微軟 Win10 Win11 專題 搜索 首頁 > 鴻蒙之家>鴻蒙手機 2026 中國電信研究院 AITMark 獎項出爐,華為小藝獲 AI 手機智能體綜合評分第一 2026/8/26 18:44:52 作者:歸瀧 責編:歸瀧 評論: 感謝網友 Autumn_Dream 的線索投遞!

剛剛
量子位AI Agent

AI視頻應用井噴,美圖打開新的增長空間

AI視頻應用迎來爆發,美圖以垂直場景和工作流編排切入市場,推出開拍、Vmake Labs及MVLAND等產品。其中開拍MAU與付費用戶大幅成長,MVLAND上線三個月內ARR翻倍,月均ARPPU達約220元,顯示美圖在視頻領域已打開新的增長空間。

剛剛
量子位AI Agent

小宇宙推出《AI趨勢報告》:AI創作、AI辦公、協作型AI等成討論新趨勢

AI內容創作者增長187%,相關節目數量增長239% 8月22日,小宇宙正式對外發布《2026 AI趨勢報告》(下稱“報告”)。 報告指出,近三年,小宇宙AI內容持續高速增長,已成為平臺增長最快、佔比最高的新品類之一。2025年,AI相關內容創作者數量同比增長187%,相關節目數量同比增長239%。 伴隨AI技術的快速爆發,小宇宙正成為AI關鍵圈層發聲、深度對話的核心聚集地。

剛剛
鈦媒體AI Agent

把500元退款交給Agent之後

企業將退款等涉及金流的權限交給AI Agent時,營運邏輯將徹底改變,必須重新設計授權鏈與責任歸屬。這不只是技術升級,而是組織流程、風險控管與稽核機制的全面重構,需確保Agent在權限內運作並保留人類介入彈性。

剛剛
VentureBeat AIAI Agent

在AI代理時代,編排成為客戶體驗的新挑戰

由Tata Communications呈現。企業部署AI代理、語音AI和自動化的速度,已超越原本支援這些技術的架構。Tata Communications客戶互動套件全球主管Gaurav Anand表示,這類部署大多是把對話式AI附加到從未為此設計的舊有系統上。Anand說:「在急於部署AI的過程中,企業大多隻是把對話式AI硬接到舊有系統上。因此,雖然許多企業採用了數位工具,但很少有企業擁有真正整合、可擴展且能無縫編排的平臺。」這個落差為人類客服人員帶來沉重的認知負擔,他們必須在零散的工具之間拼湊脈絡,才能理解AI系統的運作。

2 小時前