如何在Pandas中选择含重复列名的指定列并保留原列名
Pandas 选择重复列名的指定列(保留原列名)
问题场景
当DataFrame包含重复列名时,需要按照指定的列名列表(含重复项)选择对应列,同时保留原列名和顺序。
示例数据
原始DataFrame:
a x x x z 0 6 2 7 7 8 1 6 6 3 1 1 2 6 6 7 5 6 3 8 3 6 1 8 4 5 7 5 3 0
指定列选择列表:col_select = ["a","x","x","x"]
期望输出:
a x x x 0 6 2 7 7 1 6 6 3 1 2 6 6 7 5 3 8 3 6 1 4 5 7 5 3
现有代码问题
你当前的代码通过判断列名是否存在构建col_commun,但这种方式会把重复列名合并为唯一值,最终返回的列数远少于预期(比如这里只会返回['a','x']对应的列,而非4列)。原因是Pandas用列名字符串索引时,重复列名会返回所有同名列,但无法精确匹配你要的重复次数和顺序。
正确实现方法
核心思路是按列的位置索引选择,精准定位每个重复列的位置:
方法一:简洁迭代查找
利用df.columns.get_loc的start参数,从指定位置开始查找下一个目标列的索引:
import pandas as pd # 构造原始DataFrame data = { 'a': [6,6,6,8,5], 'x': [2,6,6,3,7], 'x': [7,3,7,6,5], 'x': [7,1,5,1,3], 'z': [8,1,6,8,0] } df = pd.DataFrame(data) col_select = ["a","x","x","x"] target_indices = [] current_pos = 0 for name in col_select: # 从current_pos开始查找下一个目标列的索引 idx = df.columns.get_loc(name, start=current_pos) target_indices.append(idx) current_pos = idx + 1 # 按位置索引选择列 df_out = df.iloc[:, target_indices] print(df_out)
方法二:列名计数映射
通过给重复列名添加计数标记,精准匹配目标列:
import pandas as pd data = { 'a': [6,6,6,8,5], 'x': [2,6,6,3,7], 'x': [7,3,7,6,5], 'x': [7,1,5,1,3], 'z': [8,1,6,8,0] } df = pd.DataFrame(data) col_select = ["a","x","x","x"] # 给原始列名添加出现次数标记 col_indices = [] name_counter = {} for col in df.columns: name_counter[col] = name_counter.get(col, 0) + 1 col_indices.append((col, name_counter[col])) # 构建目标列的位置索引 target_indices = [] target_counter = {} for name in col_select: target_counter[name] = target_counter.get(name, 0) + 1 # 找到对应(name, 次数)的列索引 for idx, (col, cnt) in enumerate(col_indices): if col == name and cnt == target_counter[name]: target_indices.append(idx) break df_out = df.iloc[:, target_indices] print(df_out)
两种方法都能准确匹配你需要的重复列,保留原列名和顺序,其中方法一更为简洁高效。
内容的提问来源于stack exchange,提问作者user21482806
相关产品推荐
相关产品推荐

