You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

特殊条件下pandas DataFrame合并及缺失值填充实现咨询

解决方案

你之前使用pd.concat()是按行堆叠两个表的所有数据,不会自动按text字段匹配关联,所以才会出现重复行和全空的对应label列。要实现按text匹配合并,应该用pandas的merge()方法做外连接。


完整实现代码

import pandas as pd

# 初始化两个示例DataFrame
df1 = pd.DataFrame({"text": ["example one", "example word one", "example two", "example sentance one"],
                    "label_1": ["O O", "O W O", "O O", "O S O"]})

df2 = pd.DataFrame({"text": ["example one", "example word one", "total example", "example sentance one"],
                    "label_2": ["O N", "O O N", "O O", "O O N"]})

# 第一步:外连接合并,按text列匹配,得到带NaN的初始结果
df = pd.merge(df1, df2, on='text', how='outer')

# 第二步(可选):将空值替换为对应单词数量的O序列
def fill_na_label(row, label_col):
    if pd.isna(row[label_col]):
        word_count = len(row['text'].split())
        return ' '.join(['O'] * word_count)
    return row[label_col]

df['label_1'] = df.apply(lambda x: fill_na_label(x, 'label_1'), axis=1)
df['label_2'] = df.apply(lambda x: fill_na_label(x, 'label_2'), axis=1)

# 输出最终结果
print(df)

运行输出结果

text label_1 label_2
0           example one     O O     O N
1      example word one   O W O   O O N
2           example two     O O     O O
3  example sentance one   O S O   O O N
4         total example     O O     O O

如果不需要替换NaN,仅保留空值,省略第二步的填充逻辑即可,第一步合并后的结果就完全匹配你给出的预期输出格式。


内容的提问来源于stack exchange,提问作者taga

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 01:36:05