Python实现提取file1中列名与file2匹配的列(支持重复列名场景)
实现方案
针对重复列名的场景,推荐直接按行原生处理,避免pandas自动重命名重复列带来的匹配问题,代码如下:
# 读取file2中指定的目标列名,去换行符后存入集合提升查找效率 with open('file2.txt', 'r', encoding='utf-8') as f: target_cols = {line.strip() for line in f if line.strip()} with open('file1.txt', 'r', encoding='utf-8') as f_in, open('file3.txt', 'w', encoding='utf-8') as f_out: # 读取表头行拆分得到原始列名 header = f_in.readline().strip().split() # 确定要保留的列索引:第一列POS固定保留,其余列名在目标集合中的全部保留 keep_indexes = [0] for idx, col in enumerate(header[1:], start=1): if col in target_cols: keep_indexes.append(idx) # 写入处理后的表头 f_out.write(' '.join([header[i] for i in keep_indexes]) + '\n') # 逐行处理剩余内容 for line in f_in: line = line.strip() if not line: continue row_items = line.split() f_out.write(' '.join([row_items[i] for i in keep_indexes]) + '\n')
如果要使用pandas实现,需要关闭pandas默认的重复列名自动重命名逻辑,代码如下:
import pandas as pd # 读取目标列名 with open('file2.txt', 'r', encoding='utf-8') as f: target_cols = {line.strip() for line in f if line.strip()} # 读取file1时关闭重复列名重命名 df = pd.read_csv('file1.txt', sep=' ', header=0, mangle_dupe_cols=False) # 筛选要保留的列,固定保留POS列 keep_cols = ['POS'] + [col for col in df.columns if col in target_cols] df_filtered = df[keep_cols] # 写入结果文件,不保留索引列 df_filtered.to_csv('file3.txt', sep=' ', index=False)
原代码问题说明
- 读取file2时没有去除每行末尾的换行符,会导致后续列名匹配失败
- pandas默认会给重复列名添加
.1、.2后缀保证列名唯一,破坏原始列名的匹配逻辑 - 筛选逻辑颠倒,不需要把df的列名加入file2的读取结果,而是要筛选file2中存在的列名对应的df列
内容的提问来源于stack exchange,提问作者Lucas
相关产品推荐
相关产品推荐

