解析保存Google Vision API响应为JSON时遇DESCRIPTOR属性错误求助
问题
测试Google Vision API手写文本识别时,能正常获取响应,但保存响应到本地时遇到AttributeError: 'RepeatedComposite' object has no attribute 'DESCRIPTOR'错误。尝试过的方法包括:
- 使用
google.protobuf.json_format中的MessageToJson和MessageToDict - 将
response.SerializeToString()传入json.loads() - 保存为二进制文件后重新加载解析JSON
- 单独保存
response.text_annotations和response.full_text_annotation
即使尝试只保存text_annotations,仍出现相同错误。当前无法运行的代码如下:
import io import os import json import pandas as pd import matplotlib.pyplot as plt import numpy as np from google.cloud import vision from google.protobuf.json_format import MessageToJson from google.protobuf import json_format vision_client = vision.ImageAnnotatorClient() path = './images/' name = 'test.jpg' with io.open(path+name, 'rb') as image_file: opened_image = image_file.read() image = vision.Image(content=opened_image) response = vision_client.document_text_detection(image=image, image_context={"language_hints": ["en-t-i0-handwrit"]}) # tried extracting only whole words here - doesn't work all_words = response.text_annotations all_words_json = MessageToJson(all_words) # Causes AttributeError: 'RepeatedComposite' object has no attribute 'DESCRIPTOR'
解决方案
错误根源
response.text_annotations是RepeatedComposite类型(重复的复合消息列表),而MessageToJson仅能处理单个protobuf消息对象,无法直接解析列表,这就是报错的原因。
方法1:直接转换完整响应为JSON
直接对整个AnnotateImageResponse对象使用MessageToJson,它能自动处理包含重复字段的完整结构:
# 转换整个响应为JSON字符串 response_json = MessageToJson(response) # 保存到本地JSON文件 with open('vision_response.json', 'w', encoding='utf-8') as f: f.write(response_json)
方法2:单独处理text_annotations列表
如果只需要保存text_annotations,需遍历列表中的每个元素,逐个转换为字典/JSON后再组合保存:
# 遍历text_annotations,逐个转换为字典 text_annotations_list = [] for annotation in response.text_annotations: annotation_dict = json_format.MessageToDict(annotation) text_annotations_list.append(annotation_dict) # 保存为JSON文件 with open('text_annotations.json', 'w', encoding='utf-8') as f: json.dump(text_annotations_list, f, indent=2, ensure_ascii=False)
方法3:二进制序列化后再转JSON
如果需要先保存二进制响应文件,后续再转为JSON,可按以下步骤处理:
# 保存响应为二进制文件 with open('response.bin', 'wb') as f: f.write(response.SerializeToString()) # 从二进制文件加载并解析为响应对象 from google.cloud.vision_v1.types import AnnotateImageResponse loaded_response = AnnotateImageResponse() with open('response.bin', 'rb') as f: loaded_response.ParseFromString(f.read()) # 转换为JSON并保存 response_json = MessageToJson(loaded_response) with open('loaded_response.json', 'w', encoding='utf-8') as f: f.write(response_json)
内容的提问来源于stack exchange,提问作者kk94
相关产品推荐
相关产品推荐

