循环合并两个DataFrame,首次循环后触发TypeError: unhashable type: 'list'
解决合并含列表类型列的DataFrame时的TypeError问题
你遇到的TypeError: unhashable type: 'list'报错,根源在于groupby('text').first()操作无法处理nlp_text列里的列表/元组这类不可哈希类型——pandas的聚合函数在处理这类类型时,会因为无法将其作为哈希键而抛出错误。结合你的需求(循环合并df_2,填充或更新df_1的nlp_text),这里有两种针对性的解决方案:
方案1:用combine_first填充NaN值(保留已有非空值)
如果你的需求是用df_2的nlp_text填充df_1中对应text的NaN值,保留已有非空的nlp_text,combine_first是最简洁的方法,它会按索引匹配自动替换NaN,完全避开groupby的问题:
import pandas as pd import numpy as np # 初始化df_1 df_1 = pd.DataFrame( {'text':['A','B','C'], 'other_info':['12','24','34'], 'nlp_text':[[('together', 'RB'),('subsidiary', 'NN')],np.NaN,np.NaN]} ).reset_index(drop=True) # 循环中生成的df_2示例 df_2 = pd.DataFrame( {'text':'B', 'nlp_text':[[('produce', 'NN'), ('sell', 'VBP')]]} # 注意这里nlp_text是元组组成的列表,要加外层括号保证每行是单个值 ) # 合并操作(循环中可重复执行) # 先将两个DataFrame的text列设为索引,方便匹配 df_1 = df_1.set_index('text') df_2 = df_2.set_index('text') # 用df_2的值填充df_1的NaN,然后恢复text为列 df_1 = df_1.combine_first(df_2).reset_index()
执行后df_1的结果:
| text | other_info | nlp_text |
|---|---|---|
| A | 12 | [('together', 'RB'), ('subsidiary', 'NN')] |
| B | 24 | [('produce', 'NN'), ('sell', 'VBP')] |
| C | 34 | NaN |
方案2:合并nlp_text列表(追加新内容)
如果你的需求是把df_2中的nlp_text追加到df_1对应text的已有列表中,可以先merge再自定义函数合并列表:
import pandas as pd import numpy as np # 初始化df_1 df_1 = pd.DataFrame( {'text':['A','B','C'], 'other_info':['12','24','34'], 'nlp_text':[[('together', 'RB'),('subsidiary', 'NN')],np.NaN,np.NaN]} ).reset_index(drop=True) # 循环中生成的df_2示例 df_2 = pd.DataFrame( {'text':'B', 'nlp_text':[[('produce', 'NN'), ('sell', 'VBP')]]} ) # 先outer merge,保留所有text merged_df = pd.merge(df_1, df_2, on='text', how='outer', suffixes=('_x', '_y')) # 定义函数合并同一text的nlp_text内容 def merge_nlp(row): combined = [] # 加入df_1中原有的非空nlp_text if not pd.isna(row['nlp_text_x']): combined.extend(row['nlp_text_x']) # 加入df_2中的非空nlp_text if not pd.isna(row['nlp_text_y']): combined.extend(row['nlp_text_y']) # 如果合并后为空则返回NaN,否则返回合并后的列表 return combined if combined else np.NaN # 应用函数生成新的nlp_text列,然后整理结构 merged_df['nlp_text'] = merged_df.apply(merge_nlp, axis=1) df_1 = merged_df[['text', 'other_info', 'nlp_text']].drop_duplicates(subset='text').reset_index(drop=True)
这种方式下,如果后续循环中再生成对应text的df_2,重复执行merge和合并步骤即可实现内容追加。
内容的提问来源于stack exchange,提问作者Kaung Myat
相关产品推荐
相关产品推荐

