MarkTechPost AI生成式AI

使用 NVIDIA srt-slurm、SLURM 配方、參數掃描與帕累託分析驗證分散式 LLM 服務基準測試

2026年7月21日 16:29

重點摘要

在本教學中,我們探討 NVIDIA 的 srt-slurm 框架,學習如何使用 srtctl 將宣告式 YAML 配置轉換為可重現的 SLURM 基準測試工作流程,用於分散式 LLM 服務。我們在 Google Colab 中設定專案,檢視其內部架構,定義叢集配置,實際執行內建與自訂配方,並為 DeepSeek-R1 建模分離式的 prefill-and-decode 部署。我們也產生參數掃描、與型別化 Python API 互動、驗證擴充配置,並透過吞吐量與延遲的帕累託前沿分析模擬基準測試結果。雖然 Colab 並未提供真實的 SLURM 環境,但我們將其作為實用的開發工作區,以理解、驗證並準備生產級別的基準測試配方。

站內 AI 整理稿

In this tutorial, we explore NVIDIA’s srt-slurm framework and learn how we use srtctl to convert declarative YAML configurations into reproducible SLURM benchmark workflows for distributed LLM serving.

We set up the project in Google Colab, inspect its internal architecture, define a cluster configuration, dry-run built-in and custom recipes, and model a disaggregated prefill-and-decode deployment for DeepSeek-R1.

We also generate parameter sweeps, interact with the typed Python API, validate expanded configurations, and analyze simulated benchmark results through a throughput-versus-latency Pareto frontier.

Although Colab does not provide a real SLURM environment, we use it as a practical development workspace to understand, validate, and prepare production-grade benchmark recipes before we submit them to an actual GPU cluster.

Copy CodeCopiedUse a different Browserimport os, sys, subprocess, textwrap, json, shutil, importlib from pathlib import Path def run(cmd, check=True, quiet=False): """Run a shell command, stream output.""" print(f"\n$ {cmd}") r = subprocess.

run(cmd, shell=True, text=True, capture_output=True) out = (r.stdout or "") + (r.stderr or "") if not quiet: print(out[-6000:]) if check and r.returncode != 0: raise RuntimeError(f"Command failed ({r.

returncode}): {cmd}") return out def section(title): print("\n" + "═"*78 + f"\n {title}\n" + "═"*78) section("1.Install srt-slurm") REPO = Path("/content/srt-slurm") if Path("/content").exists() else Path.cwd()/"srt-slurm" if not REPO.exists(): run(f"git clone --depth 1 https://github.

com/NVIDIA/srt-slurm.git {REPO}", quiet=True) run(f"{sys.executable} -m pip install -q -e {REPO}", quiet=True) sys.path.insert(0, str(REPO / "src")) importlib.invalidate_caches() os.

chdir(REPO) run("srtctl --help") We prepare the Colab environment by importing the required modules and defining reusable helper functions for command execution and section formatting.

We clone the NVIDIA srt-slurm repository, install it in editable mode, and expose its source directory to the active Python runtime.We then switch to the repository directory and verify that the srtctl command-line interface is installed correctly.Copy CodeCopiedUse a different Browsersection("2.

Repository architecture") print(textwrap.dedent(""" src/srtctl/ cli/ submit.py (apply/dry-run/preflight/monitor), do_sweep, interactive core/ schema.py (typed config), sweep.py, slurm.py (sbatch gen), validation.py, health.py, topology.py, fingerprint.py backends/ sglang.py, trtllm.py, vllm.

py, mocker.py ← engine adapters frontends/ Dynamo / router frontends templates/ Jinja2 → sbatch + orchestrator scripts recipes/ ready-made benchmarks per platform (gb200-fp4, h100, b200-fp8, qwen3-32b, dsv4-pro, mocker smoke tests, ...

) analysis/ srtlog (log parsers) + Streamlit dashboard (Pareto, latency...) docs/ sweeps.md, profiling.md, analyzing.md, config-reference.md """)) for d in ["recipes", "docs"]: print(f"{d}/ →", ", ".join(sorted(p.name for p in (REPO/d).iterdir()))[:300]) section("3.Cluster configuration (srtslurm.

yaml)") (REPO/"srtslurm.yaml").write_text(textwrap.

dedent("""\ cluster: "colab-demo" default_account: "demo-account" default_partition: "gpu" default_time_limit: "01:00:00" gpus_per_node: 4 use_gpus_per_node_directive: true use_segment_sbatch_directive: true containers: dynamo-sglang: "/containers/dynamo-sglang.sqsh" lmsysorg+sglang+v0.5.5.post2.

sqsh: "/containers/sglang-v0.5.5.sqsh" model_paths: deepseek-r1: "/models/DeepSeek-R1" """)) print((REPO/"srtslurm.yaml").read_text()) We inspect the repository structure to understand how srtctl organizes its command-line tools, schemas, backends, templates, recipes, and analysis components.

We then create a local srtslurm.yaml file containing simulated cluster defaults, container aliases, GPU settings, and model paths.We use this configuration to resolve recipe references in Colab without requiring access to an actual SLURM cluster.Copy CodeCopiedUse a different Browsersection("4.

Dry-run: mocker smoke test → generated sbatch script") run("srtctl dry-run -f recipes/mocker/agg.yaml", check=False) section("5.Custom disaggregated recipe (prefill/decode split)") (REPO/"my-disagg.yaml").write_text(textwrap.

dedent("""\ name: "colab-disagg-demo" model: path: "deepseek-r1" container: "lmsysorg+sglang+v0.5.5.post2.

sqsh" precision: "fp8" resources: gpu_type: "gb200" gpus_per_node: 4 prefill_nodes: 1 decode_nodes: 2 prefill_workers: 1 decode_workers: 2 backend: prefill_environment: { PYTHONUNBUFFERED: "1" } decode_environment: { PYTHONUNBUFFERED: "1" } sglang_config: prefill: served-model-name: "deepseek-ai/DeepSeek-R1" model-path: "/model/" trust-remote-code: true kv-cache-dtype: "fp8_e4m3" tensor-parallel-size: 4 disaggregation-mode: "prefill" decode: served-model-name: "deepseek-ai/DeepSeek-R1" model-path: "/model/" trust-remote-code: true kv-cache-dtype: "fp8_e4m3" tensor-parallel-size: 4 disaggregation-mode: "decode" benchmark: type: "sa-bench" isl: 1024 osl: 1024 concurrencies: [64, 128, 256] req_rate: "inf" """)) run("srtctl dry-run -f my-disagg.

yaml", check=False) We dry-run the built-in mocker recipe to examine how srtctl validates configurations and generates SLURM submission artifacts without executing a real benchmark.

We then define an advanced DeepSeek-R1 recipe that separates prefill and decode workloads across independent node and worker pools.We validate this disaggregated SGLang configuration through another dry run and inspect how the serving parameters are translated into job scripts.

Copy CodeCopiedUse a different Browsersection("6.Parameter sweep (grid search) — dry-run + expansion on disk") run("srtctl dry-run -f examples/example-sweep.yaml", check=False) sweep_dirs = sorted((REPO/"dry-runs").

glob("example-sweep_sweep_*")) if sweep_dirs: latest = sweep_dirs[-1] print("Per-job configs generated by the sweep expander:") for p in sorted(latest.rglob("config.yaml")): print(" ", p.relative_to(REPO)) section("7.Programmatic use of srtctl's Python API") import yaml from srtctl.core.

config import load_config from srtctl.core.sweep import generate_sweep_configs, expand_template from srtctl.core.schema import BenchmarkType, Precision, GpuType cfg = load_config("my-disagg.yaml") print(f"Loaded : {cfg.name}") print(f"Model : {cfg.model.path} ({cfg.model.precision}) on {cfg.

resources.gpu_type}") print(f"Layout : {cfg.resources.prefill_nodes}P + {cfg.resources.decode_nodes}D nodes, " f"{cfg.resources.gpus_per_node} GPUs/node") print(f"Bench : {cfg.benchmark.type} isl={cfg.benchmark.isl} osl={cfg.benchmark.osl} " f"concurrencies={cfg.benchmark.

concurrencies}") print(f"Enums : benchmarks={[b.value for b in BenchmarkType]}") print(f" precisions={[p.value for p in Precision]}, gpus={[g.value for g in GpuType]}") raw_sweep = yaml.safe_load(Path("examples/example-sweep.yaml").

read_text()) jobs = generate_sweep_configs(raw_sweep) print(f"\nSweep expands to {len(jobs)} jobs:") for job_cfg, params in jobs: pf = job_cfg["backend"]["sglang_config"]["prefill"] print(f" {params} → chunked-prefill-size={pf['chunked-prefill-size']}, " f"max-total-tokens={pf['max-total-tokens']}") print("\nTemplate substitution:", expand_template({"flag": "{x}", "n": "{y}"}, {"x": 4096, "y": 2})) We execute the example parameter sweep and inspect the individual job configurations created from its Cartesian search space.

We load our custom recipe through the typed Python API and examine its model, precision, GPU topology, benchmark settings, and supported enumeration values.We also programmatically expand sweep templates and verify how each parameter combination affects the generated backend configuration.

Copy CodeCopiedUse a different Browsersection("8.Analysis: Pareto frontier from (simulated) benchmark results") import numpy as np, matplotlib.pyplot as plt rng = np.random.default_rng(0) def simulate(variant, base_tps, base_itl): rows = [] tps_gpu = base_tps * c / (c + 90) * rng.uniform(.97, 1.

03) itl = base_itl * (1 + c/220) *

Related

相關文章

六巨頭定AI插件新標準,撞臉Claude,Anthropic沒上桌

六大科技巨頭(AWS、Anysphere、GitHub、微軟、OpenAI、Vercel)聯合發布AI智能體插件統一開放規範Agent Plugins 1.0.0,旨在統一插件打包格式,減少開發者重複勞動。該規範的結構與Anthropic的Claude Code插件系統高度相似,但Anthropic並未參與制定,而是繼續經營自己的封閉生態。

1 小時前
鈦媒體生成式AI

DeepSeek重啟融資,三年市值對齊騰訊?

DeepSeek重啟第二輪融資,以5000億元人民幣估值尋求籌集80億美元,但網傳一份由小型醫藥私募發起的專項基金募資材料引發網友質疑,後經DeepSeek員工證實部分數據屬實。該公司近期宣布API大幅漲價,可能打破其以低價換規模的估值邏輯,面臨客戶流失風險。市場關注其能否從「價格屠夫」轉型為價值提供商,以及三年內市值能否對齊騰訊等巨頭。

2 小時前

可靈AI核心技術骨幹王鑫濤被曝離職

快手可靈AI核心技術骨幹王鑫濤被曝離職,去向未知,快手官方與本人均未回應。王鑫濤是圖像與視頻生成領域知名開源項目主要作者,被視為可靈從0到1的關鍵推手。其離職發生在可靈完成獨立融資、估值180億美元的關鍵階段,可能影響研發進度與競爭優勢。

2 小時前

AI短劇、漫劇、戀綜、電影、藝人都有了,AI觀眾也不遠了

2026年AI影視內容全面爆發,從短劇、長劇到電影、綜藝,AI製作的作品大量湧現,衛視也開始播出AI短劇。AI演員如方桃子迅速走紅,商業變現能力驚人,廣告報價甚至超過許多真人網紅。AI短劇市場規模已突破220億元,用戶超過6億,但同時也引發了對真人演員就業和內容品質的擔憂。

2 小時前