You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas选取单列后concat拼接结果异常的解决求助

问题:从多列DataFrame选取单列后拼接子序列的格式异常

我有原始DataFrame og_df 和作为其子集的 sub_df,想要创建新的new_df,包含og_df中sub_df各元素对应的自身及后续n个元素。

正常示例(单列og_df场景)

import pandas as pd
og_df = pd.DataFrame({'column': range(20)})
sub_df = pd.DataFrame({'column': [1, 2, 10]})
n = 3  
new_df = pd.DataFrame({'column':[]})

for index in sub_df.index:
    new_df = pd.concat([new_df, og_df.iloc[index:index + n]])
print(new_df)

输出符合预期:

column
0        1
1        2
2        3
1        2
2        3
3        4
10      10
11      11
12      12

异常场景(多列og_df选单列操作)

当从多列og_df中用单括号[]选取单列操作时,拼接结果出现多余列和NaN:

for index in sub_df.index:
    new_df = pd.concat([new_df, og_df['column'].iloc[index:index + n]])
print(new_df)

异常输出:

column     0
1      NaN   1.0
2      NaN   2.0
3      NaN   3.0
2      NaN   2.0
3      NaN   3.0
4      NaN   4.0
10     NaN  10.0
11     NaN  11.0
12     NaN  12.0

需求:从多列og_df中仅选取一列完成上述操作,同时保证拼接结果符合预期。


解决方案

为啥会出这问题?因为用单括号og_df['column']取出来的是Series类型,而你初始化的new_df是DataFrame,concat的时候会把Series当成新列拼进去,就出现了多余的列和NaN。只要保证每次拼接的都是DataFrame类型就行,有两种简单办法:

方法1:用双括号选单列,直接拿到DataFrame

把og_df['column']改成og_df[['column']],这样选出来的是只有一列的DataFrame,和new_df结构完全匹配,拼接就不会乱:

import pandas as pd
og_df = pd.DataFrame({'column': range(20), 'other_col': range(20,40)})  # 模拟多列场景
sub_df = pd.DataFrame({'column': [1, 2, 10]})
n = 3  
new_df = pd.DataFrame({'column':[]})

for index in sub_df.index:
    new_df = pd.concat([new_df, og_df[['column']].iloc[index:index + n]])
print(new_df)

输出就是预期的结果:

column
0      1.0
1      2.0
2      3.0
1      2.0
2      3.0
3      4.0
10    10.0
11    11.0
12    12.0

方法2:把Series转成DataFrame

要是已经用单括号拿到了Series,直接调用to_frame()方法转成DataFrame再拼就行,注意指定列名和new_df保持一致:

for index in sub_df.index:
    selected_part = og_df['column'].iloc[index:index + n]
    new_df = pd.concat([new_df, selected_part.to_frame(name='column')])

这个方法也能得到完全正确的结果。

额外优化:大数据量避免循环concat

如果你的数据量很大,循环里反复concat会因为频繁创建新对象拖慢速度,不如先把所有需要的切片存到列表里,最后一次性拼接:

import pandas as pd
og_df = pd.DataFrame({'column': range(20), 'other_col': range(20,40)})
sub_df = pd.DataFrame({'column': [1, 2, 10]})
n = 3  

slice_list = []
for index in sub_df.index:
    slice_list.append(og_df[['column']].iloc[index:index + n])

new_df = pd.concat(slice_list, ignore_index=False)
print(new_df)

内容的提问来源于stack exchange,提问作者100xln2

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 13:03:36