You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于含重复列名的列表筛选Pandas DataFrame并保留重复列

问题:保留DataFrame中重复列名的子集化需求

场景说明

  • 有一个包含重复列名的列表:
    col_list = ['col1', 'col3', 'col1']
    
  • 对应的Pandas DataFrame:
    col1    col2    col3    col4
     a       c        e       g
     b       d        f       h
    
  • 期望子集化后得到包含重复列的结果(以col1, col2, col1为例):
    col1   col2     col1
       a    c        a
       b    d        b
    

尝试的代码及问题

用户尝试了以下代码:

names = ['col1', 'col4', 'col1']
df = pd.DataFrame({'col1': ['a','b'],'col2': ['c', 'd'], 'col3': ['e', 'f'], 'col4': ['g', 'h']} )
df2 = df[df.columns & names]
df2

但结果自动去重,只保留了唯一列:

col1    col4
  a      g
  b      h

解决方案

直接使用包含重复项的列表对DataFrame进行索引即可,不需要用列名交集(交集会自动去重):

示例代码

names = ['col1', 'col2', 'col1']  # 按需求定义包含重复列的列表
df = pd.DataFrame({'col1': ['a','b'],'col2': ['c', 'd'], 'col3': ['e', 'f'], 'col4': ['g', 'h']} )
df2 = df[names]
print(df2)

输出结果

col1 col2 col1
0    a    c    a
1    b    d    b

原理说明

Pandas支持DataFrame存在重复列名,直接传入包含重复项的列名列表,会严格按照列表顺序返回对应列,保留重复项。而df.columns & names会生成一个去重的列名集合,因此无法保留重复列。

内容的提问来源于stack exchange,提问作者Fatemeh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 08:52:14