You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解析保存Google Vision API响应为JSON时遇DESCRIPTOR属性错误求助

问题

测试Google Vision API手写文本识别时,能正常获取响应,但保存响应到本地时遇到AttributeError: 'RepeatedComposite' object has no attribute 'DESCRIPTOR'错误。尝试过的方法包括:

  • 使用google.protobuf.json_format中的MessageToJson和MessageToDict
  • 将response.SerializeToString()传入json.loads()
  • 保存为二进制文件后重新加载解析JSON
  • 单独保存response.text_annotations和response.full_text_annotation

即使尝试只保存text_annotations,仍出现相同错误。当前无法运行的代码如下:

import io
import os
import json
import pandas as pd
import matplotlib.pyplot as plt
import numpy as np
from google.cloud import vision
from google.protobuf.json_format import MessageToJson
from google.protobuf import json_format

vision_client = vision.ImageAnnotatorClient()
path = './images/'
name = 'test.jpg'
with io.open(path+name, 'rb') as image_file:
    opened_image = image_file.read()
image = vision.Image(content=opened_image)
response = vision_client.document_text_detection(image=image, image_context={"language_hints": ["en-t-i0-handwrit"]})

# tried extracting only whole words here - doesn't work
all_words = response.text_annotations
all_words_json = MessageToJson(all_words)

# Causes AttributeError: 'RepeatedComposite' object has no attribute 'DESCRIPTOR'

解决方案

错误根源

response.text_annotations是RepeatedComposite类型(重复的复合消息列表),而MessageToJson仅能处理单个protobuf消息对象,无法直接解析列表,这就是报错的原因。

方法1:直接转换完整响应为JSON

直接对整个AnnotateImageResponse对象使用MessageToJson,它能自动处理包含重复字段的完整结构:

# 转换整个响应为JSON字符串
response_json = MessageToJson(response)

# 保存到本地JSON文件
with open('vision_response.json', 'w', encoding='utf-8') as f:
    f.write(response_json)

方法2:单独处理text_annotations列表

如果只需要保存text_annotations,需遍历列表中的每个元素,逐个转换为字典/JSON后再组合保存:

# 遍历text_annotations,逐个转换为字典
text_annotations_list = []
for annotation in response.text_annotations:
    annotation_dict = json_format.MessageToDict(annotation)
    text_annotations_list.append(annotation_dict)

# 保存为JSON文件
with open('text_annotations.json', 'w', encoding='utf-8') as f:
    json.dump(text_annotations_list, f, indent=2, ensure_ascii=False)

方法3:二进制序列化后再转JSON

如果需要先保存二进制响应文件,后续再转为JSON,可按以下步骤处理:

# 保存响应为二进制文件
with open('response.bin', 'wb') as f:
    f.write(response.SerializeToString())

# 从二进制文件加载并解析为响应对象
from google.cloud.vision_v1.types import AnnotateImageResponse
loaded_response = AnnotateImageResponse()
with open('response.bin', 'rb') as f:
    loaded_response.ParseFromString(f.read())

# 转换为JSON并保存
response_json = MessageToJson(loaded_response)
with open('loaded_response.json', 'w', encoding='utf-8') as f:
    f.write(response_json)

内容的提问来源于stack exchange,提问作者kk94

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 06:10:39