You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现提取file1中列名与file2匹配的列(支持重复列名场景)

实现方案

针对重复列名的场景,推荐直接按行原生处理,避免pandas自动重命名重复列带来的匹配问题,代码如下:

# 读取file2中指定的目标列名,去换行符后存入集合提升查找效率
with open('file2.txt', 'r', encoding='utf-8') as f:
    target_cols = {line.strip() for line in f if line.strip()}

with open('file1.txt', 'r', encoding='utf-8') as f_in, open('file3.txt', 'w', encoding='utf-8') as f_out:
    # 读取表头行拆分得到原始列名
    header = f_in.readline().strip().split()
    # 确定要保留的列索引:第一列POS固定保留,其余列名在目标集合中的全部保留
    keep_indexes = [0]
    for idx, col in enumerate(header[1:], start=1):
        if col in target_cols:
            keep_indexes.append(idx)
    # 写入处理后的表头
    f_out.write(' '.join([header[i] for i in keep_indexes]) + '\n')
    # 逐行处理剩余内容
    for line in f_in:
        line = line.strip()
        if not line:
            continue
        row_items = line.split()
        f_out.write(' '.join([row_items[i] for i in keep_indexes]) + '\n')

如果要使用pandas实现,需要关闭pandas默认的重复列名自动重命名逻辑,代码如下:

import pandas as pd

# 读取目标列名
with open('file2.txt', 'r', encoding='utf-8') as f:
    target_cols = {line.strip() for line in f if line.strip()}

# 读取file1时关闭重复列名重命名
df = pd.read_csv('file1.txt', sep=' ', header=0, mangle_dupe_cols=False)
# 筛选要保留的列,固定保留POS列
keep_cols = ['POS'] + [col for col in df.columns if col in target_cols]
df_filtered = df[keep_cols]
# 写入结果文件,不保留索引列
df_filtered.to_csv('file3.txt', sep=' ', index=False)

原代码问题说明

  1. 读取file2时没有去除每行末尾的换行符,会导致后续列名匹配失败
  2. pandas默认会给重复列名添加.1、.2后缀保证列名唯一,破坏原始列名的匹配逻辑
  3. 筛选逻辑颠倒,不需要把df的列名加入file2的读取结果,而是要筛选file2中存在的列名对应的df列

内容的提问来源于stack exchange,提问作者Lucas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 16:36:00