如何将Google Vision API TEXT_DETECTION的JSON输出转为可读文本/Markdown?
解决方案
一、Python脚本实现文本拼接与Markdown格式化
针对Google Vision TEXT_DETECTION的碎片化输出,用Python脚本按边界框坐标分组排序,拼接成连贯文本,还能简单转成Markdown:
import json import sys from collections import defaultdict def parse_vision_text(json_data): # 按文本碎片的左上角y坐标分组,区分不同行 lines = defaultdict(list) # 跳过第一个元素(API返回的整体文本摘要) for text in json_data['responses'][0]['textAnnotations'][1:]: vertices = text['boundingPoly']['vertices'] y_pos = vertices[0]['y'] x_pos = vertices[0]['x'] lines[y_pos].append((x_pos, text['description'])) # 按行排序,每行内按x坐标排序拼接 sorted_lines = sorted(lines.items(), key=lambda item: item[0]) result = [] for _, items in sorted_lines: sorted_items = sorted(items, key=lambda i: i[0]) result.append(' '.join([item[1] for item in sorted_items])) return '\n'.join(result) def convert_to_markdown(raw_text): # 简单识别标题(短且全大写的行),可按需扩展规则 markdown_lines = [] for line in raw_text.split('\n'): stripped = line.strip() if len(stripped) < 20 and stripped.isupper(): markdown_lines.append(f'## {stripped}') else: markdown_lines.append(stripped) return '\n\n'.join(markdown_lines) if __name__ == '__main__': # 支持从管道读取JSON或传入JSON文件路径 if len(sys.argv) > 1: with open(sys.argv[1], 'r') as f: data = json.load(f) else: data = json.load(sys.stdin) raw_text = parse_vision_text(data) markdown_text = convert_to_markdown(raw_text) print(markdown_text)
使用方法
- 保存脚本为
vision_text_parser.py - 直接通过管道调用:
gcloud ml vision detect-text ./path/to/local/file.jpg | python vision_text_parser.py - 若已保存API输出的JSON文件:
python vision_text_parser.py ./vision_output.json
二、Bash组合工具快速处理
无需写脚本,用jq和awk直接在命令行处理,适合快速需求:
提取连贯文本
gcloud ml vision detect-text ./path/to/local/file.jpg | jq -r '.responses[0].textAnnotations[1:] | map([.boundingPoly.vertices[0].y, .boundingPoly.vertices[0].x, .description]) | sort_by(.[0], .[1]) | group_by(.[0]) | .[] | map(.[2]) | join(" ")'
简单转Markdown
gcloud ml vision detect-text ./path/to/local/file.jpg | jq -r '.responses[0].textAnnotations[1:] | map([.boundingPoly.vertices[0].y, .boundingPoly.vertices[0].x, .description]) | sort_by(.[0], .[1]) | group_by(.[0]) | .[] | map(.[2]) | join(" ")' | awk '{ if (length($0) < 20 && $0 ~ /^[A-Z ]+$/) print "## " $0; else print $0 }'
三、替代工具/API建议
- Google Vision DOCUMENT_TEXT_DETECTION:比TEXT_DETECTION更适合照片类文档,返回的结构自带行、段落划分,调用命令:
输出的JSON包含gcloud ml vision detect-document-text ./path/to/local/file.jpgpages.blocks.paragraphs.words层级,可直接按段落提取连贯文本,无需额外拼接。 - Rust工具:用
serde_json解析JSON并按坐标排序,适合性能敏感场景:
保存为use serde::{Deserialize}; use std::io::{self, Read}; use std::collections::BTreeMap; #[derive(Deserialize, Debug)] struct Vertex { x: i32, y: i32, } #[derive(Deserialize, Debug)] struct BoundingPoly { vertices: Vec<Vertex>, } #[derive(Deserialize, Debug)] struct TextAnnotation { description: String, bounding_poly: BoundingPoly, } #[derive(Deserialize, Debug)] struct Response { text_annotations: Vec<TextAnnotation>, } #[derive(Deserialize, Debug)] struct VisionOutput { responses: Vec<Response>, } fn main() { let mut input = String::new(); io::stdin().read_to_string(&mut input).unwrap(); let output: VisionOutput = serde_json::from_str(&input).unwrap(); let mut lines = BTreeMap::new(); for ann in &output.responses[0].text_annotations[1..] { let y = ann.bounding_poly.vertices[0].y; let x = ann.bounding_poly.vertices[0].x; lines.entry(y).or_insert_with(Vec::new).push((x, ann.description.clone())); } for (_, mut items) in lines { items.sort_by_key(|(x, _)| *x); let line: Vec<_> = items.into_iter().map(|(_, s)| s).collect(); println!("{}", line.join(" ")); } }vision_parser.rs,通过cargo build --release编译后,即可通过管道调用:gcloud ml vision detect-text ./path/to/local/file.jpg | ./target/release/vision_parser
内容的提问来源于stack exchange,提问作者Grzegorz Wierzowiecki
相关产品推荐
相关产品推荐

