运行Google Cloud自然语言API实体分析脚本遇UTF-8解析错误求助
问题描述
我尝试运行Google Cloud Natural Language API的官方Python示例脚本(完全没修改过),想对存放在Google Cloud Storage里的UTF-8文本文件gs://neotokyo-cloud-bucket/TXT/TTS-01.txt做实体分析。
在Google Cloud Shell里执行了这条命令:
python snippets.py entities-file gs://neotokyo-cloud-bucket/TXT/TTS-01.txt
结果直接弹出了protobuf的错误:
[libprotobuf ERROR google/protobuf/wire_format_lite.cc:629]. String field 'google.cloud.language.v1beta2.TextSpan.content' contains invalid UTF-8 data when parsing a protocol buffer. Use the 'bytes' type if you intend to send raw bytes.
最后返回了500内部错误:
google.api_core.exceptions.InternalServerError: 500 Exception deserializing response!
我用的示例代码是这样的:
def entities_file(gcs_uri): """Detects entities in the file located in Google Cloud Storage.""" client = language_v1beta2.LanguageServiceClient() # Instantiates a plain text document. document = types.Document( gcs_content_uri=gcs_uri, type=enums.Document.Type.PLAIN_TEXT) # Detects sentiment in the document. You can also analyze HTML with: # document.type == enums.Document.Type.HTML entities = client.analyze_entities(document).entities # entity types from enums.Entity.Type entity_type = ('UNKNOWN', 'PERSON', 'LOCATION', 'ORGANIZATION', 'EVENT', 'WORK_OF_ART', 'CONSUMER_GOOD', 'OTHER') for entity in entities: print('=' * 20) print(u'{:<16}: {}'.format('name', entity.name)) print(u'{:<16}: {}'.format('type', entity_type[entity.type])) print(u'{:<16}: {}'.format('metadata', entity.metadata)) print(u'{:<16}: {}'.format('salience', entity.salience)) print(u'{:<16}: {}'.format('wikipedia_url', entity.metadata.get('wikipedia_url', '-')))
我对protobuf完全不熟,有没有大佬能帮忙解决这个问题?
解决方案
根据错误信息来看,核心问题是API返回的响应里包含了无效的UTF-8数据,或者你的文件本身编码有问题导致API解析出错。下面是几个一步步排查修复的方法:
1. 先确认文件真的是UTF-8编码
有时候我们以为文件是UTF-8,但实际可能是带BOM的UTF-8、GBK或者ISO-8859-1这类编码。你可以在Cloud Shell里下载文件检查:
# 下载文件到本地 gsutil cp gs://neotokyo-cloud-bucket/TXT/TTS-01.txt ./TTS-01.txt # 检查编码 file -I TTS-01.txt
正常输出应该是text/plain; charset=utf-8。如果显示其他编码,比如iso-8859-1,就把文件转成标准UTF-8再传回去:
iconv -f ISO-8859-1 -t UTF-8 TTS-01.txt > TTS-01-utf8.txt gsutil cp TTS-01-utf8.txt gs://neotokyo-cloud-bucket/TXT/TTS-01.txt
2. 清理文件里的无效UTF-8字符
就算文件是UTF-8,也可能存在损坏的多字节字符(比如文件截断导致的)。用这条命令找出有问题的行:
grep -axv '.*' TTS-01.txt
找到后可以手动修复,或者用工具自动清理掉无效字符:
# -c参数会忽略无法转换的无效字符 iconv -f UTF-8 -t UTF-8 -c TTS-01.txt > TTS-01-clean.txt # 重新上传到GCS gsutil cp TTS-01-clean.txt gs://neotokyo-cloud-bucket/TXT/TTS-01.txt
3. 升级客户端库和Protobuf版本
旧版本的Google Cloud客户端库或者Protobuf可能存在解析bug,直接升级到最新版试试:
pip install --upgrade google-cloud-language protobuf
4. 切换到稳定版API(v1而非v1beta2)
你现在用的language_v1beta2是测试版API,稳定性可能不如正式版v1。把代码改成用v1版本:
# 替换原来的v1beta2导入为v1 from google.cloud import language_v1 from google.cloud.language_v1 import enums from google.cloud.language_v1 import types def entities_file(gcs_uri): """Detects entities in the file located in Google Cloud Storage.""" client = language_v1.LanguageServiceClient() # v1里参数名是type_,避免和Python关键字冲突 document = types.Document( gcs_content_uri=gcs_uri, type_=enums.Document.Type.PLAIN_TEXT) # v1的analyze_entities需要传request字典 entities = client.analyze_entities(request={'document': document}).entities entity_type = ('UNKNOWN', 'PERSON', 'LOCATION', 'ORGANIZATION', 'EVENT', 'WORK_OF_ART', 'CONSUMER_GOOD', 'OTHER') for entity in entities: print('=' * 20) print(u'{:<16}: {}'.format('name', entity.name)) # v1里实体类型的属性是type_ print(u'{:<16}: {}'.format('type', entity_type[entity.type_])) print(u'{:<16}: {}'.format('metadata', entity.metadata)) print(u'{:<16}: {}'.format('salience', entity.salience)) print(u'{:<16}: {}'.format('wikipedia_url', entity.metadata.get('wikipedia_url', '-')))
5. 绕开GCS URI,直接传递文本内容
如果上面的方法都不行,可以先把文件从GCS下载到本地,再把文本内容传给API,这样能避免API直接读取GCS文件时的编码问题:
from google.cloud import storage from google.cloud import language_v1 from google.cloud.language_v1 import enums from google.cloud.language_v1 import types def entities_file(gcs_uri): """Detects entities in the file located in Google Cloud Storage.""" # 先从GCS下载文件内容 storage_client = storage.Client() # 拆分GCS URI为桶名和文件路径 bucket_name, blob_path = gcs_uri.replace('gs://', '').split('/', 1) bucket = storage_client.bucket(bucket_name) blob = bucket.blob(blob_path) # 用UTF-8编码读取文本 text_content = blob.download_as_text(encoding='utf-8') # 创建文档并调用API client = language_v1.LanguageServiceClient() document = types.Document( content=text_content, type_=enums.Document.Type.PLAIN_TEXT) entities = client.analyze_entities(request={'document': document}).entities # 后续打印代码不变 entity_type = ('UNKNOWN', 'PERSON', 'LOCATION', 'ORGANIZATION', 'EVENT', 'WORK_OF_ART', 'CONSUMER_GOOD', 'OTHER') for entity in entities: print('=' * 20) print(u'{:<16}: {}'.format('name', entity.name)) print(u'{:<16}: {}'.format('type', entity_type[entity.type_])) print(u'{:<16}: {}'.format('metadata', entity.metadata)) print(u'{:<16}: {}'.format('salience', entity.salience)) print(u'{:<16}: {}'.format('wikipedia_url', entity.metadata.get('wikipedia_url', '-')))
内容的提问来源于stack exchange,提问作者complexitocous

