资讯动态

【拥抱鸿蒙】HarmonyOS NEXT 双路预览 + OCR:用 TaoToken 统一 Key 打通 Core Vision Kit 文字识别链路

发布时间:2026/9/30 18:15:03 来源:尧图企业网站定制
1. 双路预览 OCR 在 HarmonyOS NEXT 上到底难在哪HarmonyOS NEXT 的 Core Vision Kit 把通用文字识别封装得相当干净textRecognition.recognizeText()一个回调就能拿到结果。但真正落到「双路预览 OCR」这个场景问题往往不在识别本身而在两件事一是多路相机流怎么同时喂给 XComponent 和 ImageReceiver二是识别请求背后的 Key 与配置怎么统一管理。前者是 ArkTS 的会话配置问题后者是工程化问题。所谓双路预览就是同一颗摄像头同时输出两路流一路走 XComponent 的 Surface 做实时画面显示另一路走 ImageReceiver 拿到 JPEG 帧数据做后续处理。OCR 就挂在这第二路上对每一帧或抽帧后的 PixelMap 做文字识别。听起来顺但实际写起来createPreviewOutput要调两次、SurfaceId 要分别拿、会话beginConfig/commitConfig的顺序不能乱任何一步错了就是黑屏或者readNextImage failed。而 Key 管理这块很多教程只讲识别 API不讲识别服务怎么接。如果你用的是云端 OCR 兜底、或者想把识别结果再交给大模型做结构化提取就需要一个统一的 Key 入口。我这边习惯用 TaoToken 做统一 Key 管理一个 Key 覆盖模型对话、Coding Plan、API 调用省得在 ArkTS 工程里散落一堆密钥。官网在 https://taotoken.net/?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content API 入口是 https://taotoken.net/api 。这篇就按「先跑通双路预览 → 再挂 OCR → 再统一 Key 配置 → 最后排障」的顺序来每一步都给可复制的代码和配置。适合已经在 DevEco Studio 里建好 HarmonyOS NEXT 工程、想快速把相机文字识别跑起来的开发者。环境我实测用的是 API 11 Release、DevEco Studio 4.1.3.700真机是支持 NORMAL_PHOTO 模式的设备。先说清楚一个概念边界Core Vision Kit 的 OCR 是端侧能力不依赖网络所以它本身不需要 Key。需要 Key 的是你把识别结果往上层送的那一段比如调用大模型做语义纠错、字段抽取、多语言翻译。所以本文的 Key 配置是给「OCR 之后的链路」用的不是给recognizeText用的。这一点先分清后面配置才不会拧巴。2. TaoToken 统一 Key 前置准备与 config.toml 骨架在 ArkTS 工程里管密钥最忌讳硬编码在.ets文件里。我的做法是把 Key 和模型配置抽到一个config.toml放在entry/src/main/resources/rawfile/下运行时用resourceManager读出来。这样既方便切换环境也避免密钥进版本库时裸奔。TaoToken 的 Key 在控制台生成入口是 https://taotoken.net/console?utm_sourcetaotoken_aicg_blog_endutm_contentconsoleutm_campaignrewrite 。生成后你会拿到一个以sk-开头的字符串。注意这个 Key 是给 API 调用用的Base URL 固定为https://taotoken.net/api不要在后面乱加/v1之类的后缀具体路径以接入文档为准文档在 https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_contentdocutm_campaignrewrite 。下面是我工程里实际用的config.toml骨架路径是entry/src/main/resources/rawfile/config.toml# TaoToken 统一 Key 配置骨架 # 路径: entry/src/main/resources/rawfile/config.toml [taotoken] base_url https://taotoken.net/api api_key sk-你的Key粘贴到这里 # 默认模型OCR 后处理建议用响应快的 default_model claude-3-5-sonnet # 请求超时单位毫秒 timeout_ms 30000 [ocr] # 端侧 OCR 配置与 Key 无关但放一起便于管理 direction_detection false # 抽帧间隔避免每帧都识别导致卡顿 frame_interval_ms 800 # 送入识别的图像缩放尺寸 resize_width 320 resize_height 320 [preview] # 双路预览分辨率宽高比需与输出流一致 surface_width 1920 surface_height 1080读这个文件的 ArkTS 代码我封装成一个ConfigLoader// entry/src/main/ets/common/utils/ConfigLoader.ets import { resourceManager } from kit.LocalizationKit; import { util } from kit.ArkTS; export interface TaoTokenConfig { baseUrl: string; apiKey: string; defaultModel: string; timeoutMs: number; } export class ConfigLoader { private static instance: ConfigLoader; private config: TaoTokenConfig | undefined undefined; static getInstance(): ConfigLoader { if (!ConfigLoader.instance) { ConfigLoader.instance new ConfigLoader(); } return ConfigLoader.instance; } async load(context: Context): PromiseTaoTokenConfig { if (this.config) { return this.config; } const mgr context.resourceManager; const bytes await mgr.getRawFileContent(config.toml); const text util.TextDecoder.create(utf-8).decodeToString(bytes); this.config this.parseToml(text); return this.config; } private parseToml(text: string): TaoTokenConfig { const result: TaoTokenConfig { baseUrl: , apiKey: , defaultModel: , timeoutMs: 30000 }; const lines text.split(\n); let section ; for (const raw of lines) { const line raw.trim(); if (!line || line.startsWith(#)) { continue; } if (line.startsWith([)) { section line.replace(/[\[\]]/g, ); continue; } const idx line.indexOf(); if (idx 0) { continue; } const key line.substring(0, idx).trim(); const val line.substring(idx 1).trim().replace(/^|$/g, ); if (section taotoken) { if (key base_url) { result.baseUrl val; } else if (key api_key) { result.apiKey val; } else if (key default_model) { result.defaultModel val; } else if (key timeout_ms) { result.timeoutMs Number(val); } } } return result; } }这里没有引第三方 TOML 库是因为 ArkTS 生态里现成的 TOML 解析器不多自己写个够用的解析器反而更稳。注意getRawFileContent返回的是Uint8Array要用TextDecoder转字符串别直接toString()否则中文会乱码。Key 的权限声明别忘了在module.json5里加上网络权限{ module: { requestPermissions: [ { name: ohos.permission.INTERNET }, { name: ohos.permission.CAMERA } ] } }相机权限是动态权限首次调用cameraInput.open()前要弹窗申请这个后面排障章节会讲。到这里Key 和配置骨架就位了接下来才是重头戏——双路预览的会话配置。3. 双路预览 OCR 可复制配置与完整代码双路预览的核心是「一个 CameraInput两个 PreviewOutput」。第一路输出绑到 XComponent 的 SurfaceId负责显示第二路输出绑到 ImageReceiver 的 SurfaceId负责取帧。两路的 Profile 可以取同一个previewProfiles[0]但分辨率宽高比必须一致否则会话提交会失败。先看 XComponent 的初始化关键是onLoad里拿到 SurfaceId 并设置尺寸// entry/src/main/ets/pages/XComponentPage.ets import { camera } from kit.CameraKit; import { image } from kit.ImageKit; import { display } from kit.ArkUI; import { BusinessError } from kit.BasicServicesKit; Component export struct XComponentPage { private mXComponentController: XComponentController new XComponentController(); private xComponentSurfaceId: string ; private imageReceiver: image.ImageReceiver | undefined undefined; private cameraManager: camera.CameraManager | undefined undefined; private photoSession: camera.PhotoSession | undefined undefined; build() { Column() { XComponent({ id: LOXComponent, type: XComponentType.SURFACE, libraryname: SingleXComponent, controller: this.mXComponentController }) .onLoad(() { this.mXComponentController.setXComponentSurfaceSize({ surfaceWidth: 1920, surfaceHeight: 1080 }); this.xComponentSurfaceId this.mXComponentController.getXComponentSurfaceId(); this.initCamera(); }) .onDestroy(() { this.releaseCamera(); }) .width(100%) .height(display.getDefaultDisplaySync().width * 9 / 16) } .width(100%) .height(100%) } }libraryname: SingleXComponent对应你 native 层编译出的动态库名如果你没写 native 代码这个字段可以留空或去掉SURFACE 类型依然能工作。我实测去掉 libraryname 也能正常预览。接着是双路预览的创建这是最容易出错的地方我把顺序和注释都标清楚async createDualChannelPreview( cameraManager: camera.CameraManager, xComponentSurfaceId: string, receiver: image.ImageReceiver ): Promisevoid { // 1. 获取相机设备 const camerasDevices: camera.CameraDevice[] cameraManager.getSupportedCameras(); if (camerasDevices.length 0) { console.error(no camera device); return; } // 2. 校验是否支持拍照模式 const sceneModes: camera.SceneMode[] cameraManager.getSupportedSceneModes(camerasDevices[0]); const isSupportPhotoMode: boolean sceneModes.indexOf(camera.SceneMode.NORMAL_PHOTO) 0; if (!isSupportPhotoMode) { console.error(photo mode not support); return; } // 3. 获取输出能力 const profiles: camera.CameraOutputCapability cameraManager.getSupportedOutputCapability(camerasDevices[0], camera.SceneMode.NORMAL_PHOTO); const previewProfiles: camera.Profile[] profiles.previewProfiles; if (previewProfiles.length 0) { console.error(no preview profile); return; } // 4. 两路预览用同一个 profile保证宽高比一致 const previewProfile1: camera.Profile previewProfiles[0]; const previewProfile2: camera.Profile previewProfiles[0]; // 5. 第一路绑 XComponent Surface const previewOutput1: camera.PreviewOutput cameraManager.createPreviewOutput(previewProfile1, xComponentSurfaceId); // 6. 第二路绑 ImageReceiver Surface const receiverSurfaceId: string await receiver.getReceivingSurfaceId(); const previewOutput2: camera.PreviewOutput cameraManager.createPreviewOutput(previewProfile2, receiverSurfaceId); // 7. 创建输入并打开相机 const cameraInput: camera.CameraInput cameraManager.createCameraInput(camerasDevices[0]); await cameraInput.open(); // 8. 创建会话并配置 const photoSession: camera.PhotoSession cameraManager.createSession(camera.SceneMode.NORMAL_PHOTO) as camera.PhotoSession; photoSession.beginConfig(); photoSession.addInput(cameraInput); photoSession.addOutput(previewOutput1); photoSession.addOutput(previewOutput2); await photoSession.commitConfig(); await photoSession.start(); this.photoSession photoSession; }ImageReceiver 的创建和取帧回调initImageReceiver(): void { const receiver: image.ImageReceiver image.createImageReceiver( 1920, 1080, image.ImageFormat.JPEG, 8 ); this.imageReceiver receiver; receiver.on(imageArrival, () { receiver.readNextImage((err: BusinessError, nextImage: image.Image) { if (err || !nextImage) { console.error(readNextImage failed); return; } nextImage.getComponent(image.ComponentType.JPEG, async (err: BusinessError, imgComponent: image.Component) { if (err || !imgComponent) { console.error(getComponent failed); nextImage.release(); return; } const buffer imgComponent.byteBuffer as ArrayBuffer; if (buffer buffer.byteLength 0) { await this.ocrFromBuffer(buffer); } nextImage.release(); }); }); }); }取到 ArrayBuffer 后转 PixelMap 再识别这里尺寸要和config.toml里的resize_width/height对齐async ocrFromBuffer(buffer: ArrayBuffer): Promisevoid { const opts: image.InitializationOptions { editable: true, pixelFormat: 3, size: { height: 320, width: 320 } }; try { const pixelMap: image.PixelMap await image.createPixelMap(buffer, opts); ImageOCRUtil.recognizeText(pixelMap, (res: string) { if (res res.length 0) { console.info(OCR 识别结果: res); // 这里可以把 res 交给 TaoToken 做后处理 } }); pixelMap.release(); } catch (e) { const err e as BusinessError; console.error(createPixelMap failed: err.message); } }OCR 工具类本身沿用 Core Vision Kit 的标准写法// entry/src/main/ets/common/utils/ImageOCRUtil.ets import { textRecognition } from kit.CoreVisionKit; import { BusinessError } from kit.BasicServicesKit; import { hilog } from kit.PerformanceAnalysisKit; export class ImageOCRUtil { static async recognizeText(image: PixelMap | undefined, callback: Function): Promisevoid { if (!image) { hilog.error(0x0000, OCR, image is undefined); return; } const visionInfo: textRecognition.VisionInfo { pixelMap: image }; const config: textRecognition.TextRecognitionConfiguration { isDirectionDetectionSupported: false }; textRecognition.recognizeText(visionInfo, config, (error: BusinessError, data: textRecognition.TextRecognitionResult) { if (error.code 0) { callback(data.value.toString()); } else { hilog.error(0x0000, OCR, recognize failed: ${error.code}); } }); } }如果你要把识别结果送到 TaoToken 做结构化用http模块发请求Base URL 和 Key 从ConfigLoader取import { http } from kit.NetworkKit; async postToTaoToken(text: string, cfg: TaoTokenConfig): Promisestring { const request http.createHttp(); const body JSON.stringify({ model: cfg.defaultModel, messages: [ { role: system, content: 你是文字结构化助手把 OCR 结果整理成 JSON。 }, { role: user, content: text } ] }); const resp await request.request(cfg.baseUrl /v1/chat/completions, { method: http.RequestMethod.POST, header: { Content-Type: application/json, Authorization: Bearer cfg.apiKey }, extraData: body, connectTimeout: cfg.timeoutMs, readTimeout: cfg.timeoutMs }); request.destroy(); return resp.result.toString(); }注意Authorization头是Bearer加 Key中间一个空格别漏。路径以接入文档为准我这边用的是/v1/chat/completions这种标准形态如果你的模型走的是别的端点按文档改。4. 验证请求与成功结果从 Log 到真机画面配置写完怎么确认真的跑通了我分三层验证相机层、OCR 层、Key 链路层。相机层看 Log 里有没有photoSession.start()之后的报错。正常启动后XComponent 区域应该出现实时画面。如果黑屏先看createPreviewOutput有没有抛异常再看commitConfig是否成功。我习惯在每一步加console.info比如console.info(surfaceId1 xComponentSurfaceId); console.info(surfaceId2 receiverSurfaceId); console.info(session committed);OCR 层的验证把镜头对准一张带文字的纸观察 Log 输出。正常会看到类似OCR 识别结果: 拥抱鸿蒙 HarmonyOS NEXT OCR 识别结果: Core Vision Kit 通用文字识别如果识别结果为空字符串多半是图像尺寸或格式不对。pixelFormat: 3对应 RGBA_8888如果你拿到的 JPEG 帧解码后格式不匹配createPixelMap会失败或产出全黑图。这时候把opts里的pixelFormat换成image.PixelMapFormat.RGBA_8888试试。Key 链路层的验证单独写个测试按钮不依赖相机直接发一条文本给 TaoTokenasync testTaoToken(): Promisevoid { const cfg await ConfigLoader.getInstance().load(getContext(this)); console.info(baseUrl cfg.baseUrl); console.info(model cfg.defaultModel); const reply await this.postToTaoToken(你好返回两个字收到, cfg); console.info(TaoToken reply: reply); }成功的话 Log 里会打印出模型返回的内容。这一步能过说明 Key、Base URL、模型 ID 三件套都对。三件套缺一不可Base URL 决定请求打到哪Key 决定身份Model ID 决定用哪个模型。任何一个错了返回的要么是 401要么是模型不存在。真机上的完整验证动作我总结成四步第一步启动页面确认预览画面出现第二步把镜头对准文字确认 Log 有识别结果第三步点测试按钮确认 TaoToken 返回正常第四步把 OCR 结果和模型返回串起来确认端到端链路通。四步都过这个场景就算跑通了。有个细节要注意ImageReceiver 的readNextImage是持续回调的如果你每帧都做 OCR手机会很快发烫。我在config.toml里留了frame_interval_ms 800实际用的时候加个时间戳判断private lastOcrTime: number 0; async ocrFromBuffer(buffer: ArrayBuffer): Promisevoid { const now Date.now(); if (now - this.lastOcrTime 800) { return; } this.lastOcrTime now; // ... 后续识别逻辑 }这样既保证识别实时性又不至于把 CPU 跑满。实测下来800ms 的间隔对文字识别场景足够画面里的文字不会变化那么快。5. 本篇常见错误排查401、local proxy failed、reading choices这一节按真实报错来都是我踩过的坑。报错一401 Unauthorized或invalid api key这是 Key 链路最常见的问题。原因通常有三个Key 复制时带了空格、Authorization头格式写错、Key 已失效。先检查config.toml里api_key的值确保没有首尾空格。再检查请求头header: { Content-Type: application/json, Authorization: Bearer cfg.apiKey // Bearer 后一个空格 }如果 Key 是从控制台复制的注意别把sk-前缀漏掉。Key 失效的话去 https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_contentapi-keysutm_campaignrewrite 重新生成一个。报错二local proxy failed或连接超时这个报错通常出现在网络请求阶段说明请求根本没发出去。先确认module.json5里声明了ohos.permission.INTERNET。再确认设备网络正常。如果用的是模拟器模拟器的网络配置可能和真机不同建议直接上真机测。还有一种情况是baseUrl写错了比如多写了斜杠或者写成了https://taotoken.net/api/末尾斜杠在某些 HTTP 客户端里会导致路径拼接异常统一用https://taotoken.net/api不带尾斜杠。报错三reading choices或返回体解析失败这个报错说明请求发出去了但返回的 JSON 结构和你解析的字段对不上。常见于你把模型返回当成了固定格式去取。正确做法是先打印原始返回console.info(raw response: resp.result.toString());看清楚结构再解析。如果是流式返回resp.result可能是分片的需要按 SSE 格式逐行处理。非流式的话标准结构里内容在choices[0].message.content别取错层级。报错四readNextImage failed或getComponent failed这是相机取帧的问题和 Key 无关。原因通常是 ImageReceiver 的尺寸和预览 Profile 不匹配。createImageReceiver的宽高要和previewProfiles[0].size一致我上面写死 1920x1080如果你的设备预览 Profile 是别的分辨率这里要跟着改。另一个原因是nextImage.release()调用时机不对必须在getComponent回调处理完之后再 release提前 release 会导致 buffer 失效。报错五createPixelMap failed图像格式问题。JPEG 帧解码后的 buffer 直接喂给createPixelMap时InitializationOptions里的pixelFormat和size要和实际数据匹配。如果报错先把 buffer 的byteLength打印出来再对照宽高算一下byteLength ≈ width * height * 4RGBA。对不上就说明尺寸设错了。报错六OAuth 相关错误如果你在 TaoToken 侧用的是 OAuth 流程而不是直接 Key报错信息里会出现OAuth字样。这种情况检查 token 是否过期刷新流程是否走完。大多数端侧场景直接用 API Key 更简单OAuth 适合有后端服务的场景。排障的通用思路是先分层再定位。相机层的问题看 SurfaceId 和 ProfileOCR 层的问题看图像格式和尺寸Key 层的问题看请求头和 Base URL。三层分开测比一股脑调试快得多。接入文档在 https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_contentdocutm_campaignrewrite 遇到不确定的端点或参数先查文档再改代码。6. 把 OCR 结果接进大模型统一 Key 的长期用法双路预览加 OCR 跑通之后真正的价值在于「识别之后做什么」。纯端侧 OCR 给你的是原始文本但实际业务里往往需要结构化发票要抽金额和税号名片要抽姓名和电话文档要抽标题和正文。这些用规则写起来很累交给大模型做反而干净。这时候 TaoToken 的统一 Key 就体现出价值了。你不需要为 OCR 后处理单独接一套鉴权同一个 Key、同一个 Base URL换个 Model ID 就能从对话模型切到更适合结构化的模型。我通常把 OCR 结果拼成 prompt让模型返回 JSON然后在 ArkTS 里解析。如果你长期做这类编码和 Agent 场景可以考虑 Coding Plan入口在 https://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_contentcoding-planutm_campaignrewrite 。它适合需要持续调用、频繁切换模型的开发场景比按次调用更省心。模型对话的入口在 https://taotoken.net/chat?utm_sourcetaotoken_aicg_blog_endutm_contentchatutm_campaignrewrite 想先验证模型效果再去写代码的话可以先在那里试。回到工程实践我建议把 OCR 后处理封装成一个独立服务和相机逻辑解耦// entry/src/main/ets/common/utils/OcrPostProcessor.ets export class OcrPostProcessor { static async structure(rawText: string, cfg: TaoTokenConfig): Promisestring { const prompt 请把下面的 OCR 文本整理成 JSON字段包括 title 和 content\n${rawText}; // 复用 postToTaoToken 逻辑 return await postToTaoToken(prompt, cfg); } }这样相机页面只管取帧和识别后处理逻辑单独维护换模型、改 prompt 都不影响预览代码。Key 从ConfigLoader统一读整个工程只有一处配置改起来也方便。最后说个实际经验双路预览的功耗不低长时间开着相机加 OCR手机温度会上升。如果只是做文字识别不一定需要持续双路可以在用户点击「识别」时才启动第二路输出识别完就停。这样既省电也避免 ImageReceiver 持续回调带来的内存压力。相机会话的stop()和release()要成对调用别只 start 不 stop否则下次进页面会报设备被占用。整套链路跑下来核心就三件事会话配置别乱序、图像格式要对齐、Key 配置要统一。把这三件做扎实HarmonyOS NEXT 上的双路预览加 OCR 就能稳定跑起来。

读完文章,也想定制专属网站?

尧图设计师 24 小时内与您沟通定制方案

免费获取报价 →
↑