You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行Google Cloud自然语言API实体分析脚本遇UTF-8解析错误求助

解决Google Cloud Natural Language API实体分析时的Protobuf UTF-8解析错误

问题描述

我尝试运行Google Cloud Natural Language API的官方Python示例脚本(完全没修改过),想对存放在Google Cloud Storage里的UTF-8文本文件gs://neotokyo-cloud-bucket/TXT/TTS-01.txt做实体分析。

在Google Cloud Shell里执行了这条命令:

python snippets.py entities-file gs://neotokyo-cloud-bucket/TXT/TTS-01.txt

结果直接弹出了protobuf的错误:

[libprotobuf ERROR google/protobuf/wire_format_lite.cc:629]. String field 'google.cloud.language.v1beta2.TextSpan.content' contains invalid UTF-8 data when parsing a protocol buffer. Use the 'bytes' type if you intend to send raw bytes.

最后返回了500内部错误:

google.api_core.exceptions.InternalServerError: 500 Exception deserializing response!

我用的示例代码是这样的:

def entities_file(gcs_uri):
    """Detects entities in the file located in Google Cloud Storage."""
    client = language_v1beta2.LanguageServiceClient()

    # Instantiates a plain text document.
    document = types.Document(
        gcs_content_uri=gcs_uri,
        type=enums.Document.Type.PLAIN_TEXT)

    # Detects sentiment in the document. You can also analyze HTML with:
    # document.type == enums.Document.Type.HTML
    entities = client.analyze_entities(document).entities

    # entity types from enums.Entity.Type
    entity_type = ('UNKNOWN', 'PERSON', 'LOCATION', 'ORGANIZATION',
                   'EVENT', 'WORK_OF_ART', 'CONSUMER_GOOD', 'OTHER')

    for entity in entities:
        print('=' * 20)
        print(u'{:<16}: {}'.format('name', entity.name))
        print(u'{:<16}: {}'.format('type', entity_type[entity.type]))
        print(u'{:<16}: {}'.format('metadata', entity.metadata))
        print(u'{:<16}: {}'.format('salience', entity.salience))
        print(u'{:<16}: {}'.format('wikipedia_url', entity.metadata.get('wikipedia_url', '-')))

我对protobuf完全不熟,有没有大佬能帮忙解决这个问题?

解决方案

根据错误信息来看,核心问题是API返回的响应里包含了无效的UTF-8数据,或者你的文件本身编码有问题导致API解析出错。下面是几个一步步排查修复的方法:

1. 先确认文件真的是UTF-8编码

有时候我们以为文件是UTF-8,但实际可能是带BOM的UTF-8、GBK或者ISO-8859-1这类编码。你可以在Cloud Shell里下载文件检查:

# 下载文件到本地
gsutil cp gs://neotokyo-cloud-bucket/TXT/TTS-01.txt ./TTS-01.txt
# 检查编码
file -I TTS-01.txt

正常输出应该是text/plain; charset=utf-8。如果显示其他编码,比如iso-8859-1,就把文件转成标准UTF-8再传回去:

iconv -f ISO-8859-1 -t UTF-8 TTS-01.txt > TTS-01-utf8.txt
gsutil cp TTS-01-utf8.txt gs://neotokyo-cloud-bucket/TXT/TTS-01.txt

2. 清理文件里的无效UTF-8字符

就算文件是UTF-8,也可能存在损坏的多字节字符(比如文件截断导致的)。用这条命令找出有问题的行:

grep -axv '.*' TTS-01.txt

找到后可以手动修复,或者用工具自动清理掉无效字符:

# -c参数会忽略无法转换的无效字符
iconv -f UTF-8 -t UTF-8 -c TTS-01.txt > TTS-01-clean.txt
# 重新上传到GCS
gsutil cp TTS-01-clean.txt gs://neotokyo-cloud-bucket/TXT/TTS-01.txt

3. 升级客户端库和Protobuf版本

旧版本的Google Cloud客户端库或者Protobuf可能存在解析bug,直接升级到最新版试试:

pip install --upgrade google-cloud-language protobuf

4. 切换到稳定版API(v1而非v1beta2)

你现在用的language_v1beta2是测试版API,稳定性可能不如正式版v1。把代码改成用v1版本:

# 替换原来的v1beta2导入为v1
from google.cloud import language_v1
from google.cloud.language_v1 import enums
from google.cloud.language_v1 import types

def entities_file(gcs_uri):
    """Detects entities in the file located in Google Cloud Storage."""
    client = language_v1.LanguageServiceClient()

    # v1里参数名是type_,避免和Python关键字冲突
    document = types.Document(
        gcs_content_uri=gcs_uri,
        type_=enums.Document.Type.PLAIN_TEXT)

    # v1的analyze_entities需要传request字典
    entities = client.analyze_entities(request={'document': document}).entities

    entity_type = ('UNKNOWN', 'PERSON', 'LOCATION', 'ORGANIZATION',
                   'EVENT', 'WORK_OF_ART', 'CONSUMER_GOOD', 'OTHER')

    for entity in entities:
        print('=' * 20)
        print(u'{:<16}: {}'.format('name', entity.name))
        # v1里实体类型的属性是type_
        print(u'{:<16}: {}'.format('type', entity_type[entity.type_]))
        print(u'{:<16}: {}'.format('metadata', entity.metadata))
        print(u'{:<16}: {}'.format('salience', entity.salience))
        print(u'{:<16}: {}'.format('wikipedia_url', entity.metadata.get('wikipedia_url', '-')))

5. 绕开GCS URI,直接传递文本内容

如果上面的方法都不行,可以先把文件从GCS下载到本地,再把文本内容传给API,这样能避免API直接读取GCS文件时的编码问题:

from google.cloud import storage
from google.cloud import language_v1
from google.cloud.language_v1 import enums
from google.cloud.language_v1 import types

def entities_file(gcs_uri):
    """Detects entities in the file located in Google Cloud Storage."""
    # 先从GCS下载文件内容
    storage_client = storage.Client()
    # 拆分GCS URI为桶名和文件路径
    bucket_name, blob_path = gcs_uri.replace('gs://', '').split('/', 1)
    bucket = storage_client.bucket(bucket_name)
    blob = bucket.blob(blob_path)
    # 用UTF-8编码读取文本
    text_content = blob.download_as_text(encoding='utf-8')

    # 创建文档并调用API
    client = language_v1.LanguageServiceClient()
    document = types.Document(
        content=text_content,
        type_=enums.Document.Type.PLAIN_TEXT)

    entities = client.analyze_entities(request={'document': document}).entities

    # 后续打印代码不变
    entity_type = ('UNKNOWN', 'PERSON', 'LOCATION', 'ORGANIZATION',
                   'EVENT', 'WORK_OF_ART', 'CONSUMER_GOOD', 'OTHER')

    for entity in entities:
        print('=' * 20)
        print(u'{:<16}: {}'.format('name', entity.name))
        print(u'{:<16}: {}'.format('type', entity_type[entity.type_]))
        print(u'{:<16}: {}'.format('metadata', entity.metadata))
        print(u'{:<16}: {}'.format('salience', entity.salience))
        print(u'{:<16}: {}'.format('wikipedia_url', entity.metadata.get('wikipedia_url', '-')))

内容的提问来源于stack exchange,提问作者complexitocous

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:16:12