特殊条件下pandas DataFrame合并及缺失值填充实现咨询
解决方案
你之前使用pd.concat()是按行堆叠两个表的所有数据,不会自动按text字段匹配关联,所以才会出现重复行和全空的对应label列。要实现按text匹配合并,应该用pandas的merge()方法做外连接。
完整实现代码
import pandas as pd # 初始化两个示例DataFrame df1 = pd.DataFrame({"text": ["example one", "example word one", "example two", "example sentance one"], "label_1": ["O O", "O W O", "O O", "O S O"]}) df2 = pd.DataFrame({"text": ["example one", "example word one", "total example", "example sentance one"], "label_2": ["O N", "O O N", "O O", "O O N"]}) # 第一步:外连接合并,按text列匹配,得到带NaN的初始结果 df = pd.merge(df1, df2, on='text', how='outer') # 第二步(可选):将空值替换为对应单词数量的O序列 def fill_na_label(row, label_col): if pd.isna(row[label_col]): word_count = len(row['text'].split()) return ' '.join(['O'] * word_count) return row[label_col] df['label_1'] = df.apply(lambda x: fill_na_label(x, 'label_1'), axis=1) df['label_2'] = df.apply(lambda x: fill_na_label(x, 'label_2'), axis=1) # 输出最终结果 print(df)
运行输出结果
text label_1 label_2 0 example one O O O N 1 example word one O W O O O N 2 example two O O O O 3 example sentance one O S O O O N 4 total example O O O O
如果不需要替换NaN,仅保留空值,省略第二步的填充逻辑即可,第一步合并后的结果就完全匹配你给出的预期输出格式。
内容的提问来源于stack exchange,提问作者taga
相关产品推荐
相关产品推荐

