You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中如何将Series合并到Pandas DataFrame并消除NaN值?

解决DataFrame匹配合并后NaN值问题

你的核心问题是原代码将匹配到的两行纵向拼接(axis=0),导致每行仅包含其中一个DataFrame的列值,其余列显示NaN。正确的做法是将满足条件的两行横向合并为同一行,再添加到结果中。

修正后的循环实现

import pandas as pd
import numpy as np

# 假设IPCSection和IPCClass是已加载的DataFrame
allcolumns = np.concatenate((IPCSection.columns, IPCClass.columns), axis=0)
finalpatentclasses = pd.DataFrame(columns=allcolumns)

for _, secrow in IPCSection.iterrows():
    for _, clrow in IPCClass.iterrows():
        # 匹配第一列的包含关系
        if secrow.iloc[0] in clrow.iloc[0]:
            # 横向合并两行,转成单行DataFrame
            merged_row = pd.concat([secrow, clrow]).to_frame().T
            # 将合并后的行添加到结果
            finalpatentclasses = pd.concat([finalpatentclasses, merged_row], ignore_index=True)

display(finalpatentclasses)

更高效的无循环实现(推荐)

当数据量较大时,iterrows循环效率较低,可通过预匹配索引对的方式优化:

# 提取两个DataFrame的第一列用于匹配
sec_first_col = IPCSection.iloc[:, 0]
cls_first_col = IPCClass.iloc[:, 0]

# 收集所有满足匹配条件的索引对
matched_pairs = []
for sec_idx, sec_val in sec_first_col.items():
    matching_cls_indices = cls_first_col[cls_first_col.str.contains(sec_val)].index
    for cls_idx in matching_cls_indices:
        matched_pairs.append((sec_idx, cls_idx))

# 根据索引对提取对应行并横向合并
matched_sec_rows = IPCSection.loc[[pair[0] for pair in matched_pairs]].reset_index(drop=True)
matched_cls_rows = IPCClass.loc[[pair[1] for pair in matched_pairs]].reset_index(drop=True)

finalpatentclasses = pd.concat([matched_sec_rows, matched_cls_rows], axis=1)
display(finalpatentclasses)

关键说明

  1. 原代码中pd.concat(pdList, axis=0)会把secrow和clrow作为两行添加,导致每行仅一半列有值;
  2. 修正后使用pd.concat([secrow, clrow]).to_frame().T将两行合并为同一行,保证所有对应列都有值;
  3. 无循环实现通过批量匹配索引,避免逐行循环的性能损耗,适合处理大型数据集。

内容的提问来源于stack exchange,提问作者steliosgabriel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 15:31:41