资讯动态

Awesome Agentic AI 實戰練習 6:Function Schema 設計(bad vs good)——讓模型「挑對工具」的關鍵工程

发布时间:2026/10/9 10:02:39 来源:尧图企业网站定制
教程文档AI Agent人工智能大模型【免费下载链接】awesome-agentic-ai-zhA trilingual (繁中 / English / 简中) learning roadmap for agentic AI: from LLM basics to multi-agent systems, with 240 curated resources and hands-on examples. 中文 AI agent 學習地圖。项目地址https://gitcode.com/gh_mirrors/aw/awesome-agentic-ai-zh点击查看免费下载本指南對應 Stage 3 — 工具使用與第一個 Agent Loop 的練習 6以倉庫中examples/stage-3/06-schema-design/的starter_bad.py與starter_good.py為主體用「把攝氏 32 度換成華氏」這道同一題目對照兩種 schema 設計如何左右 LLM 的工具選擇。讀完你將掌握tool schema 為什麼是模型最依賴的 prompt、壞 schema 的典型 anti-pattern 長什麼樣、好 schema 的具體寫法型別、required、enum、結構化錯誤以及如何用 mock 測試與固定 eval 在不花錢的前提下驗證 schema 質量。完整學習方法見 docs/HOW_TO_USE.md。為什麼這題重要Schema 是 prompt 的一部分在 Agent Loop 中模型面對的不只是「問題文字」還有一份份工具的說明卡——也就是Tool Schema。Schema 會以 JSON 形式跟著 prompt 一起送進模型成為模型判斷「該呼叫哪個工具、傳什麼參數」的最主要依據。寫 schema 時不能只想「人看得懂」要想「模型能不能用它排除錯誤工具」。這題用兩個 starter 對照同一個問題「Convert 32 Celsius to Fahrenheit.」Bad schemadescription 太短、參數全部是string、沒有required、沒有enum→ LLM 容易把溫度轉換誤丟給process_dataGood schema用途明確、value: number、unit: enum[celsius, fahrenheit]、required都列好 → 用固定 eval 驗證是否較常選到convert_temperature。兩份 schema 的完整對照與可執行範例都收錄在 examples/stage-3/06-schema-design/下面會逐步深入原始碼。怎麼跑兩條路徑Path A預設、本機免費、4 個 starterPath A 使用 Ollama qwen2.5:3b完全不需要 API keyAPI 費用為$0不包含硬體、記憶體與電力成本pip install -r requirements.txt ollama pull qwen2.5:3b ollama serve python starter_bad.py # 觀察壞 schema 怎麼讓 qwen 挑錯 python starter_good.py # 觀察好 schema 怎麼讓 qwen 挑對requirements.txt 的內容只有兩行依賴非常輕量openai3.5,4 anthropic1.1,2 # Only Anthropic starters need this package.Path A 的 starter 透過openai套件連到 Ollama 的 OpenAI-compatible endpointOpenAI(base_urlhttp://localhost:11434/v1, api_keyollama)因此不需要任何真實的 API key 也能跑通完整流程。Path BAnthropic、雲端比較Path B 使用 Anthropic 的 Claude Haiku適合拿來比較不同模型對 schema 質量的反應pip install -r requirements.txt $env:ANTHROPIC_API_KEY your-key python starter_bad_anthropic.py python starter_good_anthropic.pymacOSLinux 上請改用export ANTHROPIC_API_KEYyour-key。金鑰不要寫進程式或 commit。預算方面每次先保留$0.05實際費用依輸入 tokens × $1 / 1,000,000 輸出 tokens × $5 / 1,000,000計算Tool Use 還會額外加入 prompt tokens價格查核日2026-08-27詳見 Stage 3 主文 的費用計算方式。從原始碼看兩個 Anthropic starter 的預設模型是claude-haiku-4-5-20251001並可透過環境變數覆蓋MODEL os.environ.get(MODEL, claude-haiku-4-5-20251001)見 starter_bad_anthropic.py。呼叫時設定max_tokens512從回應的contentblock 中過濾type tool_use的 block 作為工具呼叫。不花錢驗證程式邏輯mock-based兩條 path 各有一個測試檔全部使用unittest.mock取代真實 API 客戶端不打真 API、$0/runpython test.py # 驗 Path A (Ollama) starter_bad starter_good python test_anthropic.py # 驗 Path B (Anthropic) starter_*_anthropic值得注意的測試設計是每組 test 不只模擬 LLM 的選擇行為還直接檢查 schema 結構本身。以 test.py 為例def test_good_schema_has_required_fields_and_enum(): 直接檢查 schema 結構good 有 required enum、bad 沒有。 bad_temp next(t for t in bad.TOOLS_SPEC if t[function][name] convert_temperature) good_temp next(t for t in good.TOOLS_SPEC if t[function][name] convert_temperature) assert required not in bad_temp[function][parameters] assert good_temp[function][parameters][required] [value, unit] assert good_temp[function][parameters][properties][unit][enum] [celsius, fahrenheit] assert good_temp[function][parameters][additionalProperties] is False print(✅ test_good_schema_has_required_fields_and_enum)這代表「schema 質量」本身就是可被程式斷言assert的客觀屬性不依賴任何一次 LLM 的運氣。兩份測試都涵蓋五類驗證以 test_anthropic.py 對應 OpenAI-compat 版本壞 schema 可能挑錯工具mock LLM 回傳process_data驗證模糊 schema 下模型把溫度轉換誤判成「處理文字」好 schema 穩定挑對工具mock LLM 回傳convert_temperature驗證輸出{value: 89.6, unit: fahrenheit}schema 結構檢查good 有requiredenumadditionalProperties: Falsebad 沒有應用層拒絕不可信呼叫未知工具名、非法 JSON、缺欄位、多餘欄位、型別錯誤、非法 enum 值全部必須被execute_tool擋下並回傳對應錯誤訊息多個 tool call 不部分執行當模型一次回傳兩個 tool callselect_and_run會拒絕並回傳expected one tool call確保「只執行一個請求」的邊界安全。Bad vs Good schema 對照設計面向BadGoodDescriptionProcess data.Use only to summarize structured JSON table rows. Do not use for temperature conversion.參數型別全部stringnumber/array/ 對應實際型別Required無[value, unit]Enum 收斂無[celsius, fahrenheit]失敗回傳簡單字串結構化 dict retry_hint下面逐一深入原始碼看這些差異具體落在哪些欄位、以及為什麼會影響模型行為。從原始碼看 Bad schema四種 anti-pattern 一次到位starter_bad.py 的TOOLS_SPEC刻意保留了四種典型的 anti-pattern# Anti-pattern: description 太短、params 都 string、無 required、無 enum TOOLS_SPEC [ { type: function, function: { name: process_data, description: Process data., parameters: {type: object, properties: {data: {type: string}}}, }, }, { type: function, function: { name: convert_temperature, description: Convert a value., parameters: { type: object, properties: {value: {type: string}, unit: {type: string}}, }, }, }, ]逐一拆解為什麼這些寫法會害模型挑錯description 太模糊Process data.沒有說明「何時用、何時不用」。模型面對「Convert 32 Celsius to Fahrenheit」時無法從process_data的說明判斷「這不是資料處理」於是可能把溫度轉換也丟進去Convert a value.同樣沒講清楚這工具只做溫標轉換。參數全部stringvalue: string讓模型以為「32」可以是文字、unit: string讓模型可以自由填kelvin、K等任何值——型別越寬模型亂傳的自由度越高。沒有required缺required時模型可能漏填欄位而工具實作只能靠i.get(value, )給預設值硬撐產生「看似成功、其實是空值」的假結果。沒有enumunit沒有收斂成[celsius, fahrenheit]模型無法從 schema 得知合法取值範圍。對應 resources/schema-design-cheatsheet.md 中的 anti-patternAnti-2「description 是 docstring」LLM 要的是「這個 tool 什麼時候有用」不是實作細節、Anti-3「所有東西都是 string」count: string會收到five、active: string會收到yes、Anti-1「萬用工具」do_database_op把所有操作混在一個 tool 裡。從原始碼看 Good schema五條黃金規則的完整示範starter_good.py 的TOOLS_SPEC則是黃金規則的正面教材# Good schema: 用途明確、正確型別、required、enum TOOLS_SPEC [ { type: function, function: { name: process_data, description: Use only to summarize structured JSON table rows. Do not use for temperature conversion., parameters: { type: object, properties: { data: {type: array, items: {type: object}, description: Rows to inspect}, operation: {type: string, enum: [count_rows, list_columns]}, }, required: [data, operation], additionalProperties: False, }, }, }, { type: function, function: { name: convert_temperature, description: Use this when the user asks to convert temperatures between Fahrenheit and Celsius., parameters: { type: object, properties: { value: {type: number, description: Temperature value to convert}, unit: {type: string, enum: [celsius, fahrenheit], description: Unit of the input value}, }, required: [value, unit], additionalProperties: False, }, }, }, ]逐項對應 cheatsheet 的五條黃金規則規則 1description 寫給 LLM 看Use only to summarize structured JSON table rows. Do not use for temperature conversion.同時寫了情境when與排除what not負面排除直接切斷「誤用」路徑convert_temperature的 description 明確說「when the user asks to convert temperatures」。規則 2正確型別 enum 收斂value: number不是 string、data: array of object、unit與operation都用enum收斂。參照 cheatsheet 的型別收斂表unit: string→enum[celsius, fahrenheit]、count: string→integer、enabled: string→boolean、tags: string→array of string。規則 3required 分清楚required: [value, unit]、required: [data, operation]明確列出少了就不能執行的欄位。注意若把可選欄位誤列 required模型會亂編值cheatsheet 的例子把timezone列 requiredLLM 會編造「Asia/Taipei」這裡只列真正必要的欄位。規則 4tool name 參數名自說明convert_temperature(value, unit)、process_data(data, operation)動詞開頭、說清楚語意對比壞版的process_dataconvert_temperature模糊命名模型可判斷性完全不同。規則 5錯誤可恢復對應 cheatsheet 的結構化錯誤格式{error: ..., code: ..., retry_hint: ...}good schema 的工具實作回傳{error: unknown operation, retry_hint: use count_rows or list_columns}這類「含提示的失敗」讓模型有機會自我修正。好 schema 還多了一個壞 schema 沒有的欄位additionalProperties: False明確告訴模型「只能傳 schema 裡列出的欄位」從結構上堵住亂加參數的行為。應用層驗證schema 再好也不能省一個容易被忽略的教學重點即使 schema 寫得再好應用程式仍必須把模型輸出當不可信輸入重新驗證。兩個 starter 都實作了execute_tool_validate_args這道「安全應用邊界」見 starter_good.pydef execute_tool(name: str, raw_arguments: str) - tuple[dict, object]: Validate model output even when the schema is clear. if name not in TOOL_IMPL: return {}, {error: tool not allowed} try: args json.loads(raw_arguments) except (TypeError, json.JSONDecodeError): return {}, {error: arguments must be valid JSON} if not isinstance(args, dict): return {}, {error: arguments must be a JSON object} validation_error _validate_args(name, args) if validation_error: return args, {error: validation_error} try: return args, TOOL_IMPLname except (KeyError, TypeError, ValueError) as exc: return args, {error: finvalid arguments: {exc}}這道邊界做了四層檢查任何一層失敗都回傳結構化錯誤而不是崩潰allowlist 檢查工具名不在TOOL_IMPL就直接拒絕tool not allowed——這正是 Stage 3 主文 五條底線的第一條「只執行 allowlist 裡的工具不用模型輸出的名字做任意函式呼叫」JSON 語法檢查模型回傳的arguments必須是合法 JSON型別與結構檢查_validate_args欄位集合必須精確等於預期多一個、少一個都拒絕、型別必須正確value必須是int/float且排除bool、data必須是list且元素都是dict、enum 值必須合法unit只能是celsius/fahrenheit例外捕捉執行工具實作時攔截KeyError/TypeError/ValueError把異常轉成錯誤訊息。此外select_and_run還處理了兩個邊界情形沒有 tool call 時回傳observation: None多個 tool call 時拒絕並回傳{error: expected one tool call; received N}確保絕不部分執行starter_good.py。測試檔中test_multiple_calls_are_not_partially_executed專門驗證這條。兩條 SDK path 的格式差異同樣的 schema 概念在 OpenAI-compatOllama與 Anthropic 兩條 path 上欄位名略有差異這是跨供應商開發最容易踩的坑面向Path AOllamaOpenAI 格式Path BAnthropic 格式工具結構{type: function, function: {name: ..., parameters: {...}}}{name: ..., input_schema: {...}}參數 schema 欄位parametersinput_schema回應中的呼叫resp.choices[0].message.tool_calls含function.name/function.arguments字串resp.content中的tool_useblock含name/inputdictarguments 形態字串需json.loadsdict直接可用對照 starter_bad_anthropic.pyAnthropic 版的工具定義是TOOLS_SPEC [ { name: process_data, description: Process data., input_schema: {type: object, properties: {data: {type: string}}}, }, ... ]且 Anthropic 版execute_tool直接接收 dict 形態的arguments少了json.loads一步但多了一道isinstance(arguments, dict)檢查回傳arguments must be an object。兩套 starter 的 schema 語意完全一致這正是練習「同一道題、兩條 SDK path」的用意學的是 schema 設計原則本身而不是某一家 API 的細節。cheatsheet 也提醒OpenAI strict mode 是例外properties 全部要列入required、可選欄位用含null的 type 表示Anthropic、Ollama 的 strict 支援不同不要把一家規則當成通用規格。教學重點schema 品質要用固定 eval 一起測量不同 model 對 schema 質量的反應可能不同。這題的教學重點是固定 prompt、schema 與測試題用 eval 記錄行為而不是看一次成功就下結論cheatsheet 的 Anti-4。在 Ollama 上特別適合觀察這個差異觀察項Anthropic Claude haikuOllama qwen2.5:3bBad schema 是否猜對用固定 eval 測量用固定 eval 測量Good schema 是否選對用固定 eval 測量用固定 eval 測量BadGood 差距用固定 eval 測量用固定 eval 測量換句話說schema 品質與模型行為要用固定 eval 一起測量。Production 想用便宜 modelqwen / mistralschema 必須寫到能上線跑的程度。測試時要記錄的是「工具選擇、參數是否合法、程式是否拒絕未授權輸入」而不是只評最後一句話好不好看。五條黃金規則與五個 anti-pattern 速查完整的 schema 設計規則濃縮版在 resources/schema-design-cheatsheet.md核心如下5 條黃金規則description 是寫給 LLM 看的不是 docstring寫情境when與做什麼what不寫實作細節參數用對 type模糊處用 enum 收斂unit: string→enum[celsius, fahrenheit]、count: string→integer、enabled: string→boolean、tags: string→array of stringrequired vs optional 分清楚缺了就不能執行的欄位才列required有 default 不代表供應商會替你填值tool name parameter name 要自說明get_user_profile(user_id)優於fetch(id)動詞開頭、說清 querymutationactionerror 回傳要讓 LLM 可以恢復回{error: ..., code: ..., retry_hint: ...}而不是Error 500。5 個常見 anti-pattern萬用工具God Tool一個 tool 做所有事 → 拆成query_users、create_order、update_inventorydescription 是 docstring寫實作細節GET /api/v2/weather而非使用情境所有東西都是 string模型會傳five、yes、[a, b, c]只看一次成功就宣布 schema 很好固定 5–10 個正常、模糊與惡意案例對照跑沉默的失敗只回null{}會被模型當成功 → 回{success: true/false, error: ..., retry_hint: ...}。延伸挑戰與下一步跑通 starter 後可以按以下三個方向加深理解故意改壞 good schema把一個enum拿掉看 qwen 是否就開始挑錯——直接驗證「enum 收斂」到底貢獻多少選擇準確度加第三個工具寫一個跟convert_temperature用途相近但邊界模糊的 tool看 LLM 怎麼在「相近工具」之間做選擇——練習用 description 的負面排除Do not use for ...來消歧義接 examples/stage-3/05-error-handling/ 的 structured error pattern把 schema 設計 錯誤處理結合起來練習 production 級的 Agent 工具層。對照參考還包括resources/schema-design-cheatsheet.md 的詳細規則、Stage 3 主文 的完整練習序列以及 docs/HOW_TO_USE.md 的「只改一個小地方、重跑既有測試」學習方法。赞分享教程文档AI Agent人工智能大模型【免费下载链接】awesome-agentic-ai-zhA trilingual (繁中 / English / 简中) learning roadmap for agentic AI: from LLM basics to multi-agent systems, with 240 curated resources and hands-on examples. 中文 AI agent 學習地圖。项目地址https://gitcode.com/gh_mirrors/aw/awesome-agentic-ai-zh点击查看免费下载相关推荐Hello 算法資料結構章節練習題全解析知識鞏固 程式設計實戰Hello 算法資料結構章節練習題全解析知識鞏固 程式設計實戰 本文針對《Hello 算法》繁中版「資料結構」章節配套的練習題見 練習原文 http教程文档示例工程教育Flowbite 快速上手指南在 Tailwind CSS 项目中使用这套 UI 组件库与 JavaScript 交互Flowbite 快速上手指南在 Tailwind CSS 项目中使用这套 UI 组件库与 JavaScript 交互 Flowbite 是一套开源的 UI教程文档AI Agent人工智能大模型Security-101 應用程式安全實戰AppSec 關鍵能力與工具全景解析Security 101 應用程式安全實戰AppSec 關鍵能力與工具全景解析 本節Lesson 5.2是 Security 101 課程「應用程式安全基网络安全教程文档上一篇SeaTunnel 中 Avro 格式的使用基于 Kafka 连接器的读写完整指南下一篇GetQzonehistory扫码一次把QQ空间历史说说完整导出到本地创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考

读完文章,也想定制专属网站?

尧图设计师 24 小时内与您沟通定制方案

免费获取报价 →
↑