Python中展平Google实体情感分析响应对象报错咨询
解决Google实体情感分析响应对象展平Pandas DataFrame的问题
问题背景
在Python Notebook中尝试展平Pandas DataFrame(df)的entitysentiment字段,该字段存储Google实体情感分析的Protocol Buffer(protobuf)响应对象,需要遍历每行处理嵌套结构,但两次尝试的函数均报错:
- 第一个函数报错:
AttributeError: Unknown field for AnalyzeEntitySentimentResponse: split - 第二个函数报错:
AttributeError: entity_mentions
响应对象示例
entities { name: "login page" type_: OTHER salience: 0.5467509031295776 mentions { text { content: "login page" begin_offset: 24 } type_: COMMON sentiment { magnitude: 0.4000000059604645 score: -0.4000000059604645 } } sentiment { magnitude: 0.4000000059604645 score: -0.4000000059604645 } } entities { name: "app" type_: CONSUMER_GOOD salience: 0.45324909687042236 mentions { text { content: "app" begin_offset: 52 } type_: COMMON sentiment { magnitude: 0.4000000059604645 score: -0.4000000059604645 } } sentiment { magnitude: 0.4000000059604645 score: -0.4000000059604645 } } language: "en"
尝试的代码
""" # 第一个尝试的函数(报错) def extract_entities(text): entities = [] for line in text.split('\n'): if 'content:' in line: entity = line.strip().split(':')[-1].strip().replace("'", "") entities.append(entity) return entities """ # 第二个尝试的函数(报错) def extract_entities(text): entities = [] if 'entity_mentions' not in text: return entities for entity in text['entity_mentions']: entities.append(entity['content']) return entities # 应用函数到entitysentiment列 df['entity_mentions'] = df['entitysentiment'].apply(extract_entities) # 转换为单独列 entity_mentions_df = pd.DataFrame(df['entity_mentions'].to_list(), columns=['entity_mention_1', 'entity_mention_2', 'entity_mention_3']) # 合并数据框 result = pd.concat([df, entity_mentions_df], axis=1) # 删除原字段 result.drop(['entitysentiment', 'entity_mentions'], axis=1, inplace=True) # 输出结果 print(result)
错误信息
--------------------------------------------------------------------------- AttributeError Traceback (most recent call last) /tmp/ipykernel_1/3330714854.py in <module> 22 23 # Apply the function to the entitysentiment column ---> 24 df['entity_mentions'] = df['entitysentiment'].apply(extract_entities) 25 26 # Convert the entity mentions to separate columns /opt/conda/lib/python3.7/site-packages/pandas/core/series.py in apply(self, func, convert_dtype, args, **kwargs) 4355 dtype: float64 4356 """ -> 4357 return SeriesApply(self, func, convert_dtype, args, kwargs).apply() 4358 4359 def _reduce( /opt/conda/lib/python3.7/site-packages/pandas/core/apply.py in apply(self) 1041 return self.apply_str() 1042 -> 1043 return self.apply_standard() 1044 1045 def agg(self): /opt/conda/lib/python3.7/site-packages/pandas/core/apply.py in apply_standard(self) 1099 values, 1100 f, # type: ignore[arg-type] -> 1101 convert=self.convert_dtype, 1102 ) 1103 /opt/conda/lib/python3.7/site-packages/pandas/_libs/lib.pyx in pandas._libs.lib.map_infer() /tmp/ipykernel_1/3330714854.py in extract_entities(text) 13 def extract_entities(text): 14 entities = [] ---> 15 if 'entity_mentions' not in text: 16 return entities 17 for entity in text['entity_mentions']: /opt/conda/lib/python3.7/site-packages/proto/message.py in __contains__(self, key) 686 wire serialization. 687 """ -> 688 pb_value = getattr(self._pb, key) 689 try: 690 # Protocol buffers "HasField" is unfriendly; it only works AttributeError: entity_mentions
问题分析与解决思路
错误原因
- 第一个函数错误:
entitysentiment字段存储的是protobuf对象,不是字符串,无法直接调用split()方法,因此报错找不到split字段。 - 第二个函数错误:protobuf对象的字段名与代码中使用的
entity_mentions不符,从响应示例可以看到,正确的顶级字段是entities,每个实体下的提及字段是mentions,内容存储在mentions.text.content中;同时protobuf对象需要通过属性访问(.),而非字典索引([])。
正确处理函数
def extract_entities(proto_obj): entities = [] # 遍历所有实体 for entity in proto_obj.entities: # 遍历每个实体的所有提及 for mention in entity.mentions: # 提取提及内容 entities.append(mention.text.content) return entities
验证与调试技巧
如果不确定protobuf对象的结构,可以通过以下方式查看:
- 调用
dir(proto_obj)查看所有可用属性 - 调用
print(proto_obj)打印完整的对象结构,确认字段名
后续代码保留
替换函数后,原有的apply、转换列、合并数据框等步骤可以正常执行,无需修改。
内容的提问来源于stack exchange,提问作者dsx
相关产品推荐
相关产品推荐

