A Developer’s Guide to Laya: Zero-Shot Decisions and Calibration
In this tutorial, we work with Laya, the open-source decision engine from Convai Innovations that became one of the most-starred machine-learning repositories of September 2026.
Laya is a non-autoregressive System 1 model: instead of generating text, a 421-million-parameter encoder reads a piece of text and a set of typed questions, a choice between labels, a score on a scale, or a yes/no, and returns a probability for every option in a single forward pass with zero output tokens.
Its pitch is speed and calibrated probabilities, the open answer to TypeSafe’s Jev.
Rather than repeat the README’s examples, we put those promises to work on real labelled data with known answers, the banking domain of the CLINC150 intent dataset, and measure what a production router actually gets: zero-shot accuracy against a trained classifier, how much the wording and order of the options matter, how honest the shipped probabilities are, what fitting a temperature on validation data fixes and what it quietly breaks, an abstention gate fitted to an error budget, out-of-scope traffic, a yes/no question that temperature cannot repair, and typed outputs from a pydantic schema.
Copy CodeCopiedUse a different Browserimport os import sys import time import json import warnings import traceback import subprocess import urllib.
request RESULTS = {} def banner(title): print("\n" + "=" 78) print(title) print("=" 78) def section(name): def wrap(fn): def run(a, kw): banner(name) try: out = fn(a, kw) RESULTS[name] = out if isinstance(out, str) else "ok" return out except Exception as e: RESULTS[name] = f"SKIPPED / FAILED -> {type(e).
name}: {e}" print(f"\n[!] {name} did not complete: {type(e).name}: {e}") traceback.printexc(limit=3) return None return run return wrap banner("1.Install Laya and load the English checkpoint at its reviewed revision") subprocess.run([sys.executable, "-m", "pip", "install", "-q", "laya==0.3.
27"], check=True) import numpy as np import pandas as pd import torch import laya from laya.calibrate import recordsfromlabeled from laya.evals import selectiveaccuracy, aurc from laya.common import tempbucket DEVICE = "cuda" if torch.cuda.isavailable() else "cpu" # laya.
load() follows the Hub's main branch unless told otherwise.The package ships the commit # its authors reviewed for each checkpoint; pinning it keeps this notebook's weights fixed.REVISION = laya.PINNEDREVISIONS["convaiinnovations/laya"] with warnings.catchwarnings(record=True) as caught: warnings.
simplefilter("always") agent = laya.load("convaiinnovations/laya", device=DEVICE, revision=REVISION) # On CUDA Laya autocasts to fp16/bf16.Turning that off keeps every device in fp32, so a GPU run # reproduces the CPU numbers printed below to within floating-point noise.agent.
ampenabled = False SHIPPED = (list(agent.temperature), dict(agent.temperaturebyoptions)) nparams = sum(p.numel() for p in agent.model.parameters()) print(f" laya {laya.version} | torch {torch.
version} | device {DEVICE}, fp32") print(f" checkpoint convaiinnovations/laya @ {REVISION[:7]} | {nparams / 1e6:.0f}M parameters" f" | maxlen {agent.cfg['maxlen']}, headmaxlen {agent.
cfg['headmaxlen']}") print("\n Temperatures shipped with the checkpoint (probabilities = softmax(logits / T)):") for qt, name in enumerate(["choice", "score", "noul"]): print(f" {name:7s} type-level T = {SHIPPED[0][qt]:.3f}") for bucket, t in sorted(SHIPPED[1].
items()): print(f" {bucket:12s} T = {t:.3f}") for w in caught: if "temperature" in str(w.message): print("\n Warning at load time:\n " + str(w.message).replace("; ", ";\n ")) print("\n T > 1 softens probabilities and T < 1 sharpens them.Remember the choice:11+ row: the") print(" checkpoint ships 0.
10 there, which the loader clamps to 0.5, so any choice question with") print(" 11 or more options gets probabilities SHARPENED by 2x.Step 6 measures what that costs.") We install the released package, laya 0.3.27, and load the English checkpoint.Two choices here keep the run reproducible.
By default laya.load follows the Hugging Face main branch, so we pin the revision the library’s own authors reviewed, which it exposes as laya.PINNEDREVISIONS.
And on CUDA Laya autocasts to half precision, so we switch that off to keep every device in fp32 and let a GPU run reproduce the CPU numbers shown here.
Printing the checkpoint’s shipped temperatures turns up the first finding before any prediction: the entry for choice questions with eleven or more options is 0.10, outside the valid range, so the loader clamps it to 0.5 and warns.
A temperature below one sharpens probabilities, so every answer to a question with that many options will look twice as certain as the raw model is.Copy CodeCopiedUse a different BrowserTICKET = "Hi, we were billed twice for March.Please refund the duplicate today or we will cancel our plan.
" TRIAGE = { "department": {"type": "choice", "instructions": "Which department should handle this?", "criteria": {"billing": "invoices, payments, refunds", "technical": "bugs, outages, system errors", "other": "everything else"}}, "urgency": {"type": "score", "instructions": "How urgent is this?
", "criteria": ["not urgent", "soon", "blocking"]}, "churnrisk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"}, } @section("2.One forward pass, three typed questions, zero output tokens") def firstdecision(): r = agent.
predict(TICKET, TRIAGE) a = r["answers"] d, u, c = a["department"], a["urgency"], a["churnrisk"] print(f" state: {TICKET!r}\n") print(f" department (choice) -> {d['choice']!r} probabilities {d['probabilities']}") print(f" urgency (score) -> {u['score']:.2f} on 0..
2 level probabilities {u['probabilities']}") print(f" churnrisk (noul) -> P(yes) = {c['noul']:.3f}") print("\n Two confidence fields, two different quantities:") for qid, ans in a.items(): print(f" {qid:11s} answerconfidence {ans['answerconfidence']:.3f} confidence {ans['confidence']:.
3f}") print(" answerconfidence is the probability of the reported answer, max(p): the number that") print(" calibration, the abstention gate and every metric below use.confidence is 1 - normalised") print(" entropy, whose scale depends on the number of options.Do not threshold on it.
") print(f"\n usage: {r['usage']}") print(" outputtokens is always 0: Laya scores the options it is given and never generates text.") return f"{d['choice']} / urgency {u['score']:.2f} / P(churn) {c['noul']:.
2f} in one pass" firstdecision() One call to predict answers three typed questions about a support ticket in a single forward pass: the department as a choice, the urgency as a score from 0 to 2, and the churn risk as a yes/no.
The result carries a probability for every option and two confidence fields that are easy to confuse.answerconfidence is the probability of the reported answer, and it is the quantity that calibration, the abstention gate and every metric later in this tutorial use.
confidence is one minus the normalized entropy, whose scale depends on how many options a question has.The usage block shows zero output tokens, because Laya scores the options it is given and never generates text.Copy CodeCopiedUse a different Browser@section("3.
What a pass costs: questions are rows, options are nearly free") def costmodel(): def medianms(q, n=7): agent.predict(TICKET, q) times = [] for in range(n): t0 = time.perfcounter() r = agent.predict(TICKET, q) times.append(1000 * (time.perfcounter() - t0)) return float(np.
median(times)), r["usage"]["inputtokens"] print(f" {'one state, asking ...':34s} {'ms (median of 7)':>16s} {'inputtokens':>13s}") rows = {} for n in (1, 4, 16): q = {f"q{i}": {"type": "noul", "instructions": f"Does the message mention topic number {i}?
"} for i in range(n)} rows[f"{n} yes/no"] = medianms(q) print(f" {f'{n:2d} yes/no questions':34s} {rows[f'{n} yes/no'][0]:16.1f} {rows[f'{n} yes/no'][1]:13d}") for k in (3, 15, 40): q = {"pick": {"type": "choice", "in
Related
相關文章

AI算力硬合作,馬斯克還是更相信中國製造
馬斯克確認正與臺積電討論Terafab芯片製造合作,此前英特爾是唯一被點名的夥伴,計劃提供14A製程。Terafab規模龐大,目標年產1TW AI算力,約為當前全球產出50倍,馬斯克可能藉此分散風險並確保芯片供應。外界推測合作可能採臺積電直接運營、SpaceX控股或僅技術合作等模式,英特爾的14A製程也可能與臺積電技術混搭使用。
NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction
Back to Articles NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction Enterprise + Article Published September 29, 2026 Upvote - Jingang Qu jingangqu Follow nvidia Valter Hudovernik valterh-nv Follow nvidia Martin Jurkovic martinjurkovic Follow nvidia Cedric Lorenz cedric-lorenz Follow nvidia Akihiro Nitta akihironitta-nv Follow nvidia DG dgordeev Follow nvidia Federico Lopez flopeznv Follow nvidia Ramona Bendias RBendiasN Follow nvidia Gilberto Titericz Jr titericz Follow nvidia Aleksandar S.
黃仁勳:明年英偉達芯片銷量將翻倍,AI監管不該照搬社交媒體模式
NVIDIA創辦人黃仁勳預估明年該公司晶片銷售量將翻倍成長,反映市場對運算力的強勁需求。他同時認為AI監管不應直接套用社交媒體模式,以免扼殺創新,呼籲決策者深入了解技術特性。

美國電動汽車企業 Lucid 計劃在歐洲部署至少 2.5 萬輛 Robotaxi
Lucid 的合作方為歐洲共享出行平臺 Bolt。該 Robotaxi 將基於 Lucid 的即將推出的中型車平臺,採用 NVIDIA Hyperion 開發平臺和參考體系架構。

每小時“燒錢”超20萬元,羅福莉直播小米大模型訓練
小米大模型MiMo-V2.6正處於強化學習中間階段,每秒需調用數千張高階GPU運算,每小時成本高達數十萬元。業界指出強化學習階段的訓練資源消耗極為驚人,每小時「燒錢」超過20萬元已是普遍共識。

華為昇騰960超節點落地 AI算力競爭從單芯片轉向系統工程
華為正式發表昇騰960超節點運算架構,標誌AI算力競爭從單晶片轉向系統工程。該系統透過高速互聯與分布式計算最佳化,解決大規模訓練的延遲與調度問題,滿足千億參數模型需求。此舉反映華為在晶片供應受限下,以系統層級整合提升整體效率,拓展差異化競爭優勢。