You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Google Vision API TEXT_DETECTION的JSON输出转为可读文本/Markdown?

解决方案

一、Python脚本实现文本拼接与Markdown格式化

针对Google Vision TEXT_DETECTION的碎片化输出,用Python脚本按边界框坐标分组排序,拼接成连贯文本,还能简单转成Markdown:

import json
import sys
from collections import defaultdict

def parse_vision_text(json_data):
    # 按文本碎片的左上角y坐标分组,区分不同行
    lines = defaultdict(list)
    # 跳过第一个元素(API返回的整体文本摘要)
    for text in json_data['responses'][0]['textAnnotations'][1:]:
        vertices = text['boundingPoly']['vertices']
        y_pos = vertices[0]['y']
        x_pos = vertices[0]['x']
        lines[y_pos].append((x_pos, text['description']))
    
    # 按行排序,每行内按x坐标排序拼接
    sorted_lines = sorted(lines.items(), key=lambda item: item[0])
    result = []
    for _, items in sorted_lines:
        sorted_items = sorted(items, key=lambda i: i[0])
        result.append(' '.join([item[1] for item in sorted_items]))
    
    return '\n'.join(result)

def convert_to_markdown(raw_text):
    # 简单识别标题(短且全大写的行),可按需扩展规则
    markdown_lines = []
    for line in raw_text.split('\n'):
        stripped = line.strip()
        if len(stripped) < 20 and stripped.isupper():
            markdown_lines.append(f'## {stripped}')
        else:
            markdown_lines.append(stripped)
    return '\n\n'.join(markdown_lines)

if __name__ == '__main__':
    # 支持从管道读取JSON或传入JSON文件路径
    if len(sys.argv) > 1:
        with open(sys.argv[1], 'r') as f:
            data = json.load(f)
    else:
        data = json.load(sys.stdin)
    
    raw_text = parse_vision_text(data)
    markdown_text = convert_to_markdown(raw_text)
    
    print(markdown_text)

使用方法

  1. 保存脚本为vision_text_parser.py
  2. 直接通过管道调用:
    gcloud ml vision detect-text ./path/to/local/file.jpg | python vision_text_parser.py
    
  3. 若已保存API输出的JSON文件:
    python vision_text_parser.py ./vision_output.json
    

二、Bash组合工具快速处理

无需写脚本,用jq和awk直接在命令行处理,适合快速需求:

提取连贯文本

gcloud ml vision detect-text ./path/to/local/file.jpg | jq -r '.responses[0].textAnnotations[1:] | map([.boundingPoly.vertices[0].y, .boundingPoly.vertices[0].x, .description]) | sort_by(.[0], .[1]) | group_by(.[0]) | .[] | map(.[2]) | join(" ")'

简单转Markdown

gcloud ml vision detect-text ./path/to/local/file.jpg | jq -r '.responses[0].textAnnotations[1:] | map([.boundingPoly.vertices[0].y, .boundingPoly.vertices[0].x, .description]) | sort_by(.[0], .[1]) | group_by(.[0]) | .[] | map(.[2]) | join(" ")' | awk '{ if (length($0) < 20 && $0 ~ /^[A-Z ]+$/) print "## " $0; else print $0 }'

三、替代工具/API建议

  • Google Vision DOCUMENT_TEXT_DETECTION:比TEXT_DETECTION更适合照片类文档,返回的结构自带行、段落划分,调用命令:
    gcloud ml vision detect-document-text ./path/to/local/file.jpg
    
    输出的JSON包含pages.blocks.paragraphs.words层级,可直接按段落提取连贯文本,无需额外拼接。
  • Rust工具:用serde_json解析JSON并按坐标排序,适合性能敏感场景:
    use serde::{Deserialize};
    use std::io::{self, Read};
    use std::collections::BTreeMap;
    
    #[derive(Deserialize, Debug)]
    struct Vertex {
        x: i32,
        y: i32,
    }
    
    #[derive(Deserialize, Debug)]
    struct BoundingPoly {
        vertices: Vec<Vertex>,
    }
    
    #[derive(Deserialize, Debug)]
    struct TextAnnotation {
        description: String,
        bounding_poly: BoundingPoly,
    }
    
    #[derive(Deserialize, Debug)]
    struct Response {
        text_annotations: Vec<TextAnnotation>,
    }
    
    #[derive(Deserialize, Debug)]
    struct VisionOutput {
        responses: Vec<Response>,
    }
    
    fn main() {
        let mut input = String::new();
        io::stdin().read_to_string(&mut input).unwrap();
        let output: VisionOutput = serde_json::from_str(&input).unwrap();
        
        let mut lines = BTreeMap::new();
        for ann in &output.responses[0].text_annotations[1..] {
            let y = ann.bounding_poly.vertices[0].y;
            let x = ann.bounding_poly.vertices[0].x;
            lines.entry(y).or_insert_with(Vec::new).push((x, ann.description.clone()));
        }
        
        for (_, mut items) in lines {
            items.sort_by_key(|(x, _)| *x);
            let line: Vec<_> = items.into_iter().map(|(_, s)| s).collect();
            println!("{}", line.join(" "));
        }
    }
    
    保存为vision_parser.rs,通过cargo build --release编译后,即可通过管道调用:
    gcloud ml vision detect-text ./path/to/local/file.jpg | ./target/release/vision_parser
    

内容的提问来源于stack exchange,提问作者Grzegorz Wierzowiecki

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 08:07:02