1. 从一次“假完成”说起Runtime 判定任务完成的信号机制到底是什么如果你正在做 Agent 开发大概率遇到过这种场景模型信誓旦旦地说“任务已完成”结果你去检查文件代码根本没改或者它陷入“继续搜索 → 继续搜索”的死循环跑了 40 轮还在原地打转。这类问题在社区里有个很形象的说法叫“假完成”。它本质上不是模型能力问题而是 Runtime 没有一套可靠的完成判定链路。Runtime 判定任务完成靠的不是单一信号而是三类信号的交叉验证LLM 输出的终止符、Skill 返回的结构化状态、以及 Context 的收敛条件。LLM 输出终止符是最表层的一层比如模型吐出Task Complete、DONE、或者工具调用列表为空。但这一层最不可靠因为模型可能在信息不足时“礼貌性收尾”。Skill 返回状态是第二层比如文件写入 Skill 返回{status:ok,bytes_written:1024}测试 Skill 返回{exit_code:0}。Context 收敛是第三层判断最近 N 轮是否还有新的 Observation 注入、Token 增长是否趋于平缓、目标条件是否被满足。我试过在一个代码修复 Agent 里只依赖第一层信号结果 10 次里有 3 次是假完成。后来把三层信号做成“与门”逻辑假完成率降到接近零。这套判定链路要落地需要一个统一的请求通道来观测每次 LLM 调用的返回结构TaoToken 的 API 通道正好适合做这件事——它把模型对话、Skill 调用、Context 状态都收敛到同一个 Key 下日志埋点不用在多个供应商之间来回切换。这篇文章会带你从零搭一套可观测的完成判定链路先讲清楚 Runtime 的事件循环里完成信号从哪来再给出可复制的回调配置和日志埋点最后用 TaoToken 统一 Key 发起请求通过返回结构验证任务是否真正结束。适合正在写 Agent Runtime、或者被“假完成”和死循环折磨的开发者。2. TaoToken 前置准备统一 Key 通道与 Runtime 观测的关系在动手写判定逻辑之前先把请求通道理顺。Agent Runtime 在运行过程中会频繁调用 LLM如果每次调用都散落在不同的供应商、不同的 Key、不同的返回格式上完成判定就无从谈起。TaoToken 在这里扮演的角色是统一入口你用一个 Key 就能访问多种模型返回结构保持一致Runtime 的日志埋点只需要解析一种格式。先拿到 Key。打开官网 https://taotoken.net/?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content 注册后在控制台创建 API Key。控制台地址是 https://taotoken.net/console?utm_sourcetaotoken_aicg_blog_endutm_contentconsoleutm_campaignrewrite Key 管理页面在 https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_contentapi-keysutm_campaignrewrite 。创建时建议给 Key 起一个能标识用途的名字比如agent-runtime-prod方便后面在日志里区分。API 的基础地址是 https://taotoken.net/api 注意这个地址不带 UTM 参数直接用于代码里的base_url。如果你用的是 OpenAI 兼容的 SDK把base_url指向它即可。模型 ID 可以在模型对话页面 https://taotoken.net/chat?utm_sourcetaotoken_aicg_blog_endutm_contentmodel-chatutm_campaignrewrite 里先试跑一下确认返回结构符合预期再写进 Runtime。这里要强调一个容易被忽略的点Runtime 的完成判定依赖返回结构里的finish_reason、tool_calls、usage三个字段。不同供应商对这三个字段的填充方式有差异有的把finish_reason填成stop有的填成tool_calls有的在流式返回里根本不带。TaoToken 的统一通道会把这些字段规范化Runtime 只需要按一套逻辑解析。这就是为什么建议在接入阶段就用统一 Key而不是等 Runtime 写完再回头适配。如果你打算长期跑编码类 Agent可以了解一下 Coding Plan https://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_contentcoding-planutm_campaignrewrite 它针对高频编码场景做了额度优化。接入文档在 https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_contentdocutm_campaignrewrite 里面有完整的请求示例和字段说明写 Runtime 回调时对着看会省很多时间。3. 可复制配置Runtime 回调、日志埋点与 settings 片段这一节给出可以直接抄的配置。先看 Runtime 的回调结构。下面是一个 Python 版的完成判定回调核心逻辑是三层信号与门LLM 终止符、Skill 状态、Context 收敛。# runtime_completion.py import json import time from dataclasses import dataclass, field from typing import Any dataclass class CompletionSignal: llm_terminated: bool False skill_ok: bool False context_converged: bool False reason: str trace: dict field(default_factorydict) class RuntimeCompletionChecker: def __init__(self, converge_window: int 3, token_growth_threshold: float 0.05): self.converge_window converge_window self.token_growth_threshold token_growth_threshold self.recent_observations [] self.token_history [] def on_llm_response(self, resp: dict) - CompletionSignal: sig CompletionSignal() choice resp.get(choices, [{}])[0] finish choice.get(finish_reason, ) content choice.get(message, {}).get(content, ) or tool_calls choice.get(message, {}).get(tool_calls, []) or [] # 第一层LLM 终止符 if finish stop and not tool_calls: sig.llm_terminated True if any(kw in content for kw in [Task Complete, DONE, 任务完成]): sig.llm_terminated True # 记录 token 增长 usage resp.get(usage, {}) total usage.get(total_tokens, 0) self.token_history.append(total) if len(self.token_history) self.converge_window: self.token_history.pop(0) sig.trace[finish_reason] finish sig.trace[tool_call_count] len(tool_calls) sig.trace[total_tokens] total return sig def on_skill_result(self, skill_name: str, result: dict) - bool: # 第二层Skill 返回状态 status result.get(status, ) exit_code result.get(exit_code, None) ok status in (ok, success) and (exit_code is None or exit_code 0) self.recent_observations.append({skill: skill_name, ok: ok, ts: time.time()}) if len(self.recent_observations) self.converge_window: self.recent_observations.pop(0) return ok def check_context_convergence(self) - bool: # 第三层Context 收敛 if len(self.token_history) self.converge_window: return False growth (self.token_history[-1] - self.token_history[0]) / max(self.token_history[0], 1) no_new_obs all(not o[ok] for o in self.recent_observations[-self.converge_window:]) return growth self.token_growth_threshold and not no_new_obs def decide(self, llm_sig: CompletionSignal, skill_ok: bool) - CompletionSignal: llm_sig.skill_ok skill_ok llm_sig.context_converged self.check_context_convergence() if llm_sig.llm_terminated and llm_sig.skill_ok and llm_sig.context_converged: llm_sig.reason all_signals_pass elif llm_sig.llm_terminated and not llm_sig.skill_ok: llm_sig.reason fake_completion_skill_failed elif not llm_sig.llm_terminated and llm_sig.context_converged: llm_sig.reason possible_dead_loop else: llm_sig.reason continue return llm_sig再看日志埋点配置。用一个 JSON 文件把埋点规则固定下来Runtime 每次判定都写一条结构化日志方便后面用jq或日志平台检索。{ runtime_logging: { version: 1.0, sink: stdout, format: json_lines, fields: { trace_id: string, step: int, finish_reason: string, tool_call_count: int, total_tokens: int, skill_name: string, skill_status: string, context_converged: bool, decision: string, reason: string }, sampling: { always_log_reasons: [fake_completion_skill_failed, possible_dead_loop], sample_rate_normal: 0.2 } } }如果你用的是 Claude Code 或类似的 Agent 工具配置通常放在settings.json里。下面是一个接入片段把 Base URL、Key、Model ID 三件套写全同时打开 Runtime 的完成判定日志。{ env: { ANTHROPIC_BASE_URL: https://taotoken.net/api, ANTHROPIC_API_KEY: sk-your-taotoken-key, ANTHROPIC_MODEL: claude-sonnet-4-20250514 }, runtime: { completion_check: { enabled: true, converge_window: 3, token_growth_threshold: 0.05, log_sink: stdout } } }注意ANTHROPIC_BASE_URL后面不要带斜杠也不要带 UTM 参数直接写 https://taotoken.net/api 即可。Key 从环境变量注入不要硬编码在文件里。Model ID 按你实际使用的模型填写可以在模型对话页面确认。4. 验证请求用 TaoToken 统一 Key 发起调用并检查返回结构配置写好后先跑一个最小验证请求确认返回结构里包含完成判定需要的字段。下面这段代码用 OpenAI 兼容 SDK 发起一次对话打印finish_reason、tool_calls、usage三个关键字段。# verify_completion.py import os from openai import OpenAI client OpenAI( base_urlhttps://taotoken.net/api, api_keyos.environ[TAOTOKEN_API_KEY], ) resp client.chat.completions.create( modelclaude-sonnet-4-20250514, messages[ {role: system, content: 你是一个任务执行 Agent。完成任务后输出 Task Complete。}, {role: user, content: 把 hello.txt 的内容改成 world然后确认。}, ], temperature0, ) choice resp.choices[0] print(finish_reason:, choice.finish_reason) print(content:, choice.message.content) print(tool_calls:, choice.message.tool_calls) print(usage:, resp.usage)跑完之后你会看到类似这样的输出finish_reason: stop content: 已修改 hello.txt内容为 world。Task Complete tool_calls: None usage: CompletionUsage(completion_tokens42, prompt_tokens128, total_tokens170)这里finish_reason是stoptool_calls为空content里包含Task Complete第一层信号通过。但注意这次请求里没有真实的 Skill 调用所以第二层信号是缺失的。要验证完整链路需要把 Skill 执行结果也接进来。下面是一个带 Skill 调用的验证脚本模拟 Runtime 执行文件写入 Skill 后把结果喂给完成判定器。# verify_full_chain.py import json from runtime_completion import RuntimeCompletionChecker checker RuntimeCompletionChecker(converge_window3) # 模拟三轮 LLM 响应 responses [ {choices: [{finish_reason: tool_calls, message: {content: , tool_calls: [{id: 1, function: {name: write_file}}]}}], usage: {total_tokens: 150}}, {choices: [{finish_reason: tool_calls, message: {content: , tool_calls: [{id: 2, function: {name: run_test}}]}}], usage: {total_tokens: 180}}, {choices: [{finish_reason: stop, message: {content: Task Complete, tool_calls: []}}], usage: {total_tokens: 185}}, ] skill_results [ (write_file, {status: ok, bytes_written: 5}), (run_test, {status: ok, exit_code: 0}), ] for i, resp in enumerate(responses): sig checker.on_llm_response(resp) if i len(skill_results): name, result skill_results[i] ok checker.on_skill_result(name, result) else: ok True final checker.decide(sig, ok) print(json.dumps({ step: i, finish_reason: sig.trace[finish_reason], total_tokens: sig.trace[total_tokens], skill_ok: ok, context_converged: final.context_converged, decision: final.reason, }, ensure_asciiFalse))预期输出{step: 0, finish_reason: tool_calls, total_tokens: 150, skill_ok: true, context_converged: false, decision: continue} {step: 1, finish_reason: tool_calls, total_tokens: 180, skill_ok: true, context_converged: false, decision: continue} {step: 2, finish_reason: stop, total_tokens: 185, skill_ok: true, context_converged: true, decision: all_signals_pass}到第三步三层信号全部通过decision是all_signals_passRuntime 可以安全停止循环并返回最终答案。如果第二步的 Skill 返回exit_code: 1那么第三步的decision会变成fake_completion_skill_failedRuntime 就知道模型说完成了但实际没完成应该继续修复而不是直接返回。这套验证跑通后你可以把RuntimeCompletionChecker挂到真实的事件循环里。每次 LLM 返回后调on_llm_response每次 Skill 执行后调on_skill_result循环末尾调decide。日志按第 3 节的 JSON 格式输出fake_completion_skill_failed和possible_dead_loop两种 reason 建议全量记录方便事后复盘。5. 本篇常见错排查401、local proxy failed、reading choices、OAuth接入过程中最容易撞上的几类报错这里逐个对照。401 Unauthorized。最常见的原因是 Key 没注入到环境变量或者 Key 前面多了空格。检查echo $TAOTOKEN_API_KEY是否输出完整字符串。另一个原因是把 Key 写在了base_url里比如https://sk-xxxtaotoken.net/api这种写法不被支持。正确做法是base_url只写 https://taotoken.net/api Key 通过api_key参数或Authorization: Bearer头传递。如果确认 Key 没问题还是 401去 API Keys 页面 https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_contentapi-keysutm_campaignrewrite 确认 Key 状态是否正常、额度是否耗尽。local proxy failed。这个报错通常出现在 Runtime 配置了本地代理但代理进程没起来或者代理端口被占用。先检查 Runtime 的settings.json里有没有残留的proxy字段如果有且你不需要代理直接删掉。如果确实需要走本地代理确认代理进程在监听并且base_url指向的是代理地址而不是 https://taotoken.net/api 。注意这里说的代理是本地开发环境的 HTTP 代理不是网络访问层面的东西两者不要混淆。reading choices 报错。典型信息是KeyError: choices或list index out of range。原因是返回结构里没有choices字段或者choices是空列表。这通常发生在请求被拒绝但 HTTP 状态码是 200 的情况下返回体里是一个错误对象而不是正常的 completion 结构。排查方法是在解析前先打印完整resp确认resp.choices存在且非空。如果返回体里有error字段按错误信息处理。另一个可能是流式返回没处理完就解析确保流式场景下等[DONE]之后再取choices。OAuth 相关报错。如果你用的是 Claude Code 或 Codex 这类工具它们可能默认走 OAuth 登录而不是 API Key。报错信息里会出现OAuth token expired或invalid_grant。解决方式是在工具的配置里显式指定 API Key 模式把ANTHROPIC_API_KEY或OPENAI_API_KEY设好同时把ANTHROPIC_BASE_URL指向 https://taotoken.net/api 。如果工具同时支持 OAuth 和 API Key确认配置优先级API Key 配置通常要放在更靠前的位置。Codex 的auth.json里如果残留了 OAuth 字段建议清空后只保留 API Key 相关配置。还有一个隐蔽的坑finish_reason是length而不是stop。这意味着模型输出被 max_tokens 截断了任务可能没真正完成。Runtime 的完成判定里要把length视为“未完成”继续下一轮而不是当成终止符。这个 case 在长任务里很常见日志里看到finish_reason: length就要检查 max_tokens 设置。6. 把完成判定链路接进你的 Runtime到这里三层信号与门的判定链路已经完整了。LLM 终止符负责第一层过滤Skill 状态负责第二层校验Context 收敛负责第三层兜底。三者都通过才判定任务真正完成任何一层不通过都继续循环或触发修复。实际落地时建议把RuntimeCompletionChecker做成单例每个任务一个实例trace_id贯穿整个任务生命周期。日志按 JSON Lines 输出fake_completion_skill_failed和possible_dead_loop两种 reason 全量记录正常完成按 20% 采样。这样既不会日志爆炸又能在出问题时快速定位。如果你还在选 Runtime 的请求通道TaoToken 的统一 Key 方案值得先试一下。模型对话页面 https://taotoken.net/chat?utm_sourcetaotoken_aicg_blog_endutm_contentmodel-chatutm_campaignrewrite 可以直接验证返回结构接入文档 https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_contentdocutm_campaignrewrite 里有完整的字段说明。长期跑编码类 Agent 的话Coding Plan https://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_contentcoding-planutm_campaignrewrite 在额度上更划算。最后留一个实用技巧在 Runtime 的循环里加一个硬性步数上限比如 50 步。即使三层信号都没触发到 50 步也强制退出并记录possible_dead_loop。这个兜底能防止极端情况下的无限循环代价只是偶尔多跑几步。