You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中合并数据框行:添加行并填充缺失值NA

解决方案

直接使用pandas的concat合并两个DataFrame,缺少的列会自动填充NA,再通过临时标记列控制排序顺序,确保同一Sample下原有数据在前、新增数据在后:

步骤1:导入模块并构造示例数据

import pandas as pd

# 构造原始大数据框df1
df1 = pd.DataFrame({
    'Sample': ['patient1', 'patient1', 'patient2', 'patient3', 'patient3'],
    'Gene': ['Gene1', 'Gene2', 'Gene2', 'Gene1', 'Gene3'],
    'P.Value': [0.05, 0.05, 0.05, 0.05, 0.05],
    'Q value': [0.03, 0.02, 0.03, 0.03, 0.03]
})

# 构造待添加的df2
df2 = pd.DataFrame({
    'Sample': ['patient1', 'patient1', 'patient2', 'patient3'],
    'Gene': ['gene_a', 'gene_b', 'gene_b', 'gene_a']
})

步骤2:合并并排序

# 添加临时标记列,区分原始数据(0)和新增数据(1)
df1['_source'] = 0
df2['_source'] = 1

# 合并两个DataFrame,自动为df2缺少的列填充NA
combined_df = pd.concat([df1, df2], ignore_index=True)

# 按Sample分组,每组内先显示原始数据,再显示新增数据
combined_df = combined_df.sort_values(by=['Sample', '_source'])

# 移除临时标记列,重置索引
combined_df = combined_df.drop(columns='_source').reset_index(drop=True)

输出结果

执行后得到的combined_df即为你期望的格式:

Sample   Gene  P.Value  Q value
0  patient1  Gene1     0.05     0.03
1  patient1  Gene2     0.05     0.02
2  patient1  gene_a      NaN      NaN
3  patient1  gene_b      NaN      NaN
4  patient2  Gene2     0.05     0.03
5  patient2  gene_b      NaN      NaN
6  patient3  Gene1     0.05     0.03
7  patient3  Gene3     0.05     0.03
8  patient3  gene_a      NaN      NaN

这个方法完全适配你150列、900行的数据集,处理效率不会有问题。

内容的提问来源于stack exchange,提问作者Aryh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 11:17:24