Google Cloud Vision API自3月7日起误识别#旁“23”字符问题求助
Google Cloud Vision API 识别#符号时误生成不存在的“23”字符问题
- 问题现象:自3月7日起,Google Cloud Vision API会在#符号的左侧或右侧识别出不存在的“23”字符。例如图片中实际内容为“HIGHWAY # 6 NORTH”,却被识别成“HIGHWAY # 236 NORTH”。
- 复现情况:通过google-cloud-vision Python API返回的对象可确认该问题,旧版本及最新3.7.2版本均能复现。API返回的相关符号数据示例如下:
... symbols { bounding_box { vertices { x: 166 y: 7 } vertices { x: 178 y: 7 } vertices { x: 178 y: 28 } vertices { x: 166 y: 28 } } text: "#" confidence: 0.431674063 } confidence: 0.431674063 } words { property { detected_languages { language_code: "en" confidence: 1 } } bounding_box { vertices { x: 165 y: 7 } vertices { x: 201 y: 7 } vertices { x: 201 y: 28 } vertices { x: 165 y: 28 } } symbols { bounding_box { vertices { x: 165 y: 7 } vertices { x: 178 y: 7 } vertices { x: 178 y: 28 } vertices { x: 165 y: 28 } } text: "2" confidence: 0.822502255 } symbols { bounding_box { vertices { x: 165 y: 7 } vertices { x: 183 y: 7 } vertices { x: 183 y: 28 } vertices { x: 165 y: 28 } } text: "3" confidence: 0.882505536 } ...
- 推测原因:曾猜测该问题与#的Unicode编码U+0023有关,但此前从未出现过类似问题。结合发布说明,Google于12月5日后90天(即3月5日)切换了新OCR模型,推测该问题与此模型更新相关。
请问是否有其他用户遇到过相同的问题?
内容的提问来源于stack exchange,提问作者Travis Lu
相关产品推荐
相关产品推荐

