You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中展平Google实体情感分析响应对象报错咨询

解决Google实体情感分析响应对象展平Pandas DataFrame的问题

问题背景

在Python Notebook中尝试展平Pandas DataFrame(df)的entitysentiment字段,该字段存储Google实体情感分析的Protocol Buffer(protobuf)响应对象,需要遍历每行处理嵌套结构,但两次尝试的函数均报错:

  • 第一个函数报错:AttributeError: Unknown field for AnalyzeEntitySentimentResponse: split
  • 第二个函数报错:AttributeError: entity_mentions

响应对象示例

entities {
  name: "login page"
  type_: OTHER
  salience: 0.5467509031295776
  mentions {
    text {
      content: "login page"
      begin_offset: 24
    }
    type_: COMMON
    sentiment {
      magnitude: 0.4000000059604645
      score: -0.4000000059604645
    }
  }
  sentiment {
    magnitude: 0.4000000059604645
    score: -0.4000000059604645
  }
}
entities {
  name: "app"
  type_: CONSUMER_GOOD
  salience: 0.45324909687042236
  mentions {
    text {
      content: "app"
      begin_offset: 52
    }
    type_: COMMON
    sentiment {
      magnitude: 0.4000000059604645
      score: -0.4000000059604645
    }
  }
  sentiment {
    magnitude: 0.4000000059604645
    score: -0.4000000059604645
  }
}
language: "en"

尝试的代码

"""
# 第一个尝试的函数(报错)
def extract_entities(text):
    entities = []
    for line in text.split('\n'):
        if 'content:' in line:
            entity = line.strip().split(':')[-1].strip().replace("'", "")
            entities.append(entity)
    return entities
"""

# 第二个尝试的函数(报错)
def extract_entities(text):
    entities = []
    if 'entity_mentions' not in text:
        return entities
    for entity in text['entity_mentions']:
        entities.append(entity['content'])
    return entities

# 应用函数到entitysentiment列
df['entity_mentions'] = df['entitysentiment'].apply(extract_entities)

# 转换为单独列
entity_mentions_df = pd.DataFrame(df['entity_mentions'].to_list(), columns=['entity_mention_1', 'entity_mention_2', 'entity_mention_3'])

# 合并数据框
result = pd.concat([df, entity_mentions_df], axis=1)

# 删除原字段
result.drop(['entitysentiment', 'entity_mentions'], axis=1, inplace=True)

# 输出结果
print(result)

错误信息

---------------------------------------------------------------------------
AttributeError                            Traceback (most recent call last)
/tmp/ipykernel_1/3330714854.py in <module>
     22 
     23 # Apply the function to the entitysentiment column
---> 24 df['entity_mentions'] = df['entitysentiment'].apply(extract_entities)
     25 
     26 # Convert the entity mentions to separate columns

/opt/conda/lib/python3.7/site-packages/pandas/core/series.py in apply(self, func, convert_dtype, args, **kwargs)
   4355         dtype: float64
   4356         """
-> 4357         return SeriesApply(self, func, convert_dtype, args, kwargs).apply()
   4358 
   4359     def _reduce(

/opt/conda/lib/python3.7/site-packages/pandas/core/apply.py in apply(self)
   1041             return self.apply_str()
   1042 
-> 1043         return self.apply_standard()
   1044 
   1045     def agg(self):

/opt/conda/lib/python3.7/site-packages/pandas/core/apply.py in apply_standard(self)
   1099                     values,
   1100                     f,  # type: ignore[arg-type]
-> 1101                     convert=self.convert_dtype,
   1102                 )
   1103 

/opt/conda/lib/python3.7/site-packages/pandas/_libs/lib.pyx in pandas._libs.lib.map_infer()

/tmp/ipykernel_1/3330714854.py in extract_entities(text)
     13 def extract_entities(text):
     14     entities = []
---> 15     if 'entity_mentions' not in text:
     16         return entities
     17     for entity in text['entity_mentions']:

/opt/conda/lib/python3.7/site-packages/proto/message.py in __contains__(self, key)
    686                 wire serialization.
    687         """
-> 688         pb_value = getattr(self._pb, key)
    689         try:
    690             # Protocol buffers "HasField" is unfriendly; it only works

AttributeError: entity_mentions

问题分析与解决思路

错误原因

  1. 第一个函数错误:entitysentiment字段存储的是protobuf对象,不是字符串,无法直接调用split()方法,因此报错找不到split字段。
  2. 第二个函数错误:protobuf对象的字段名与代码中使用的entity_mentions不符,从响应示例可以看到,正确的顶级字段是entities,每个实体下的提及字段是mentions,内容存储在mentions.text.content中;同时protobuf对象需要通过属性访问(.),而非字典索引([])。

正确处理函数

def extract_entities(proto_obj):
    entities = []
    # 遍历所有实体
    for entity in proto_obj.entities:
        # 遍历每个实体的所有提及
        for mention in entity.mentions:
            # 提取提及内容
            entities.append(mention.text.content)
    return entities

验证与调试技巧

如果不确定protobuf对象的结构,可以通过以下方式查看:

  • 调用dir(proto_obj)查看所有可用属性
  • 调用print(proto_obj)打印完整的对象结构,确认字段名

后续代码保留

替换函数后,原有的apply、转换列、合并数据框等步骤可以正常执行,无需修改。

内容的提问来源于stack exchange,提问作者dsx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 17:54:59