资讯动态

Elasticsearch AI Indices 实战:让 agents 用 ES|QL 保留答案,无需通读内容

发布时间:2026/9/29 20:33:39 来源:尧图企业网站定制
1. 为什么 agents 不该通读全文做过 RAG 的人大概都遇到过这个场景用户问一句「《Vikings》里演 Torvi 的女演员还因为什么出名」agent 转头就去browsecomp-plus这种全文索引里捞文档一次MATCH(text, ...)拉回十条每条正文几千字全塞进上下文。结果就是输入 token 直接冲到 38 万延迟 40 多秒工具调用 8 次其中还有两次read_file在翻分页结果。答案是对的但代价高得离谱而且每次检索失败都会再叠一层成本。Elasticsearch AI Indices 想解决的就是这件事。它的核心思路是把「事实」提前算好存成一个结构化、可混合检索的记录agent 用一次 ES|QL 查询就能拿到答案而不是把整篇文档读进上下文。这个预计算出来的上下文单元官方叫 Knowledge Indicator简称 KI。KI 是什么可以把它理解成文档的「事实卡片」标题、摘要、能回答哪些问题、关键实体、主题标签全部基于原文、不编造。它存在一个特殊的 Elasticsearch index 里index 名以ai-index-开头字段用semantic_text开箱即用地支持混合搜索。适合谁适合已经在用 RAG、但被 token 成本和延迟卡住的团队也适合想让 agent 在多个 harnessKibana、Claude Code、LangChain Deep Agents之间复用同一套检索逻辑的人。我试过把同一套 KI 接到不同 agent 上检索逻辑几乎不用改因为 skill 本质就是「指令 一条 ES|QL」。下面按可跟做的顺序把索引映射、workflow、ES|QL 查询和端到端验证走一遍。2. TaoToken 前置给 agent 准备一个稳定的模型入口KI 的生成和 agent 的推理都离不开 LLM。这一步不是 Elasticsearch 的必需项但如果你想让 workflow 里的ai.agent步骤和本地 Deep Agents 脚本都能稳定调用模型建议先把模型入口配好。TaoToken 提供 OpenAI 兼容的接口官网在 https://taotoken.net/?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content API 基址是 https://taotoken.net/api 。它的作用是让你用一套base_urlapi_key就能在 Kibana workflow、Python 脚本、Claude Code 之间切换模型不用每个 harness 单独改配置。具体操作分三步。先去控制台创建密钥地址是 https://taotoken.net/console?utm_sourcetaotoken_aicg_blog_endutm_contentconsoleutm_campaignrewrite 在 API Keys 页面生成一个 key地址是 https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_contentapi-keysutm_campaignrewrite 。拿到 key 后在本地脚本里这样设置环境变量export LLM_BASE_URLhttps://taotoken.net/api export LLM_API_KEYsk-你的key export LLM_MODELanthropic/claude-sonnet-4.5如果你更习惯在对话界面里先验证模型是否通可以直接打开模型对话页 https://taotoken.net/chat?utm_sourcetaotoken_aicg_blog_endutm_contentmodel-chatutm_campaignrewrite 发一条测试消息。长期跑编码类 agent、需要固定额度和并发的话可以看 Coding Plan https://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_contentcoding-planutm_campaignrewrite 。接入细节和参数说明在文档里 https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_contentdocutm_campaignrewrite Claude Code 相关的配置在 https://taotoken.net/claudecode-anthropic?utm_sourcetaotoken_aicg_blog_endutm_contentclaudecodeutm_campaignrewrite 。注意LLM_BASE_URL末尾不要带/v1OpenAI 兼容客户端会自己拼路径。如果报 404先检查这一项。3. 可复制配置索引映射、AI Index 与 KI 生成 workflow3.1 建源语料索引 browsecomp-plus先准备数据源。这里用 BrowseComp-Plus 语料做示例它是一批人工校验过的网页文档适合做推理型检索基准。创建索引时把字段和描述写清楚后面 agent 看 mapping 时能少猜PUT browsecomp-plus { mappings: { _meta: { description: BrowseComp-Plus corpus: human-verified web documents. BM25-only index. }, properties: { docid: { type: keyword }, text: { type: text }, title: { type: text }, url: { type: keyword } } } }数据用_bulkAPI 灌进去字段就是docid / title / url / text。示例里只挑了几篇文档因为给全量语料每篇都生成 KI 会跑很久跟做练习没必要。3.2 建 AI Index 存 KIAI Index 是专门存 KI 的索引命名以ai-index-开头创建时已经预置好必需 mappingsemantic_text直接可用PUT ai-index-idx-my-corpus这一步不用手写字段开箱即用的混合搜索能力已经就位。KI 文档里我们会写入type: corpus_entry、title、content、description、tags、attributes等字段其中content和description是语义检索的主要面。3.3 用 Kibana Workflow 把文档提炼成 KIWorkflow 的逻辑是一次 ES|QL 读出目标文档逐个用ai.agent生成结构化 KI再bulk写回 AI Index。关键点是用docid作为_id保证重复运行是幂等的upsert 而不是重复插入。version: 1 name: browsecomp-plus-doc-ki description: Query corpus with ES|QL, generate a KI per doc, bulk-write into AI Index. enabled: true triggers: - type: manual steps: - name: query_corpus type: elasticsearch.esql.query with: query: FROM browsecomp-plus | WHERE text IS NOT NULL AND docid IN (11589,50639,64501,41758,57766,84983,82008) | KEEP docid, title, url, text | EVAL text SUBSTRING(text, 1, 12000) - name: loop_corpus_docs type: foreach foreach: {{ steps.query_corpus.output.values }} steps: - name: generate_ki type: ai.agent timeout: 300s with: message: You are a knowledge engineer building a Knowledge Indicator (KI). Be 100% grounded: never state anything not supported by the text. Prefer concrete named specifics over vague phrasing. Document ID: {{ foreach.item[0] }} Title: {{ foreach.item[1] }} URL: {{ foreach.item[2] }} Body: {{ foreach.item[3] }} schema: type: object properties: title: { type: string } summary: { type: string } answers_questions: { type: array, items: { type: string } } key_entities: { type: array, items: { type: string } } topics: { type: array, items: { type: string } } tagline: { type: string } required: [title, summary, answers_questions, key_entities, topics] - name: sink_ki type: elasticsearch.bulk with: index: ai-index-idx-my-corpus operations: - index: _id: {{ foreach.item[0] }} - timestamp: {{ execution.startedAt | date: %Y-%m-%dT%H:%M:%S.%LZ }} type: corpus_entry title: {{ steps.generate_ki.output.structured_output.title }} tags: [browsecomp-plus] attributes: docid: {{ foreach.item[0] }} url: {{ foreach.item[2] }} source_index: browsecomp-plus content: KNOWLEDGE INDICATOR {{ steps.generate_ki.output.structured_output.summary }} description: {{ steps.generate_ki.output.structured_output.tagline }}几个容易踩的点SUBSTRING(text, 1, 12000)是为了把 prompt 控制在上下文窗口内不截断的话整篇正文会撑爆foreach.item[N]的下标由KEEP的列顺序决定item[0]docid、item[1]title、item[2]url、item[3]text顺序错了字段就串了sink_ki里显式写_id是幂等的关键。3.4 写一个 query-ki skillskill 就是给 agent 的检索指令跟具体 harness 无关。存成skills/query-ki/SKILL.md--- name: query-ki description: Retrieve Knowledge Indicators from the Elasticsearch AI Index before answering. allowed-tools: esql_query --- # Retrieving Knowledge Indicators KIs live in indices named ai-index-*. Call esql_query with the query below. Substitute the users question for query, and corpus_entry as ki_type. FROM ai-index-idx-* METADATA _id, _index, _score | WHERE type ki_type | FORK (WHERE MATCH(content, query) OR MATCH(description, query) | SORT _score DESC | LIMIT 20) (WHERE MATCH(content.semantic, query) OR MATCH(description.semantic, query) | SORT _score DESC | LIMIT 20) | FUSE | SORT _score DESC | KEEP title, content, description, tags | LIMIT 5 Ground your answer in what the query returns, and cite the KI titles you used. If nothing relevant comes back, say so rather than guessing.这条 ES|QL 做了三件事按type corpus_entry过滤出事实类 KI用FORK同时跑 BM25 和语义两路检索用FUSE做 reciprocal rank fusion 融合默认就是 RRF不用额外配参数。4. 验证请求一次端到端跑通并对比结果4.1 先看 KI 有没有写进去Workflow 跑完后直接查 AI Index 确认FROM ai-index-idx-* | WHERE type corpus_entry | KEEP title, description, attributes, tags | LIMIT 25你应该能看到类似Vikings (TV series) - Wikipedia这样的 KIattributes.docid是57766description里带着 tagline 和 topics。这说明预计算这一步成功了。4.2 基线 agent直接查全文索引先建一个不带 skill 的 agent让它直接查browsecomp-plus测出基线成本。核心工具就一个esql_queryfrom elasticsearch import Elasticsearch from langchain_core.tools import tool from langchain_openai import ChatOpenAI from deepagents import create_deep_agent import os, sys, time es Elasticsearch(os.environ[ES_URL], api_keyos.environ[ES_API_KEY]) tool def esql_query(query: str): Execute an ES|QL query. Full-text syntax: WHERE MATCH(field, value). resp es.esql.query(queryquery, formatjson) cols [c[name] for c in resp[columns]] return [dict(zip(cols, row)) for row in resp[values]] baseline_agent create_deep_agent( modelChatOpenAI( base_urlos.environ.get(LLM_BASE_URL, https://taotoken.net/api), modelos.environ.get(LLM_MODEL, anthropic/claude-sonnet-4.5), api_keyos.environ[LLM_API_KEY], ), tools[esql_query], system_prompt( Answer by querying the raw index browsecomp-plus with ES|QL. Full-text syntax: WHERE MATCH(field, \value\). Ground your answer strictly in the rows returned, and cite the docid. ), ) start time.perf_counter() result baseline_agent.invoke({messages: [{role: user, content: sys.argv[1]}]}) print(fLatency: {time.perf_counter() - start:.2f}s) print(result[messages][-1].content)用问题「What was the actress who played Torvi from Vikings also known for?」跑基线实测输出是工具调用 8 次6 次esql_query全打在browsecomp-plus上2 次read_file翻分页token 386,187输入 384,940 / 输出 1,247延迟 44.86 秒。答案正确但输入 token 几乎全花在把整篇正文读进上下文上。4.3 KI agent走 query-ki skill再建一个带 skill 的 agent工具只留esql_query通过FilesystemBackend从磁盘加载skills目录from deepagents.backends.filesystem import FilesystemBackend backend FilesystemBackend(root_dir., virtual_modeFalse) agent create_deep_agent( modelChatOpenAI( base_urlos.environ.get(LLM_BASE_URL, https://taotoken.net/api), modelos.environ.get(LLM_MODEL, anthropic/claude-sonnet-4.5), api_keyos.environ[LLM_API_KEY], ), tools[esql_query], skills[skills], backendbackend, system_prompt( When a question depends on specific facts, names, dates, or events, use the query-ki skill to retrieve Knowledge Indicators before answering. Ground your answer strictly in what it returns, and cite the KI titles. ), )同一个问题KI agent 的输出是工具调用 2 次1 次read_file读 SKILL.md1 次esql_query打在ai-index-idx-*上token 27,625输入 27,037 / 输出 588延迟 15.22 秒。答案同样正确同样有依据引用了Georgia Hirst和Georgia Hirst - Wikipedia两个 KI。4.4 对比结果指标基线无 AI Index使用 AI Index工具调用总数82read_file 调用21esql_query 调用6全打 browsecomp-plus1打 ai-index-idx-*消耗 token386,18727,625延迟44.86 秒15.22 秒答案有依据、正确有依据、正确token 降了约 93%延迟降了约三分之二答案质量没变。具体数字会随运行和 agent 不同浮动但量级差异是稳定的。5. 本篇常见错排查ES|QL 语法写错全文检索必须写成WHERE MATCH(field, value)不能写成field MATCH value。这是最常见的报错来源agent 也经常在这里翻车建议在 system prompt 里显式写清楚。foreach.item[N]字段串位下标由KEEP的列顺序决定。如果你把KEEP docid, title, url, text改成别的顺序item[0]就不再是 docid写进 KI 的字段会全乱。改查询时一定同步改下标。重复运行产生重复 KIsink_ki里必须显式设置_id: {{ foreach.item[0] }}否则每次跑 workflow 都会新增一批文档。设了_id之后是 upsert幂等。prompt 超上下文窗口不截断正文直接喂给ai.agent长文档会撑爆窗口导致报错。用SUBSTRING(text, 1, 12000)限制长度或者在生产环境改用ai.prompt降低成本。skill 没被加载FilesystemBackend(root_dir., virtual_modeFalse)的root_dir要指向包含skills目录的位置skills[skills]是相对root_dir的路径。路径不对的话 agent 根本读不到 SKILL.md会退化成直接查全文索引。AI Index 查不到结果先确认type过滤值对得上。workflow 写入时用的是corpus_entry查询时WHERE type corpus_entry必须一致大小写和拼写都要对。模型入口 404检查LLM_BASE_URL是否误加了/v1以及 key 是否有效。可以在模型对话页先发一条消息确认通路。6. 把 KI 接到你的 agent 工作流里到这里一条完整的链路就跑通了源文档进browsecomp-plusworkflow 把每篇提炼成 KI 写进ai-index-idx-my-corpusagent 通过 query-ki skill 用一次 ES|QL 混合检索拿到事实直接作答全程不读全文。如果你卡在接入或排障上先去 API Keys 页面 https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_contentapi-keysutm_campaignrewrite 确认密钥再对照接入文档 https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_contentdocutm_campaignrewrite 检查base_url和参数。想先验证模型是否正常用模型对话页 https://taotoken.net/chat?utm_sourcetaotoken_aicg_blog_endutm_contentmodel-chatutm_campaignrewrite 发一条消息最快。如果你要把这套 KI 检索长期挂在编码或 Agent 流程里跑Coding Plan https://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_contentcoding-planutm_campaignrewrite 能提供更稳定的额度。最后留一个实操建议KI 的 schema 不要一次定死。先按summary / answers_questions / key_entities / topics跑通观察 agent 实际引用哪些字段再决定要不要加dates、figures这类更细的字段。字段越多生成成本越高检索面也未必更准。

读完文章,也想定制专属网站?

尧图设计师 24 小时内与您沟通定制方案

免费获取报价 →
↑