保留Pandas DataFrame中列名包含列表中子串的列
解决DataFrame按子串筛选列并合并的问题
我来帮你搞定这个问题!你现在的循环思路是对的,但需要调整一下方式来合并筛选结果,这里有几个实用的方法:
方法一:先收集所有符合条件的列名(推荐)
这种方法不用反复生成小DataFrame,效率更高,步骤很清晰:
- 先把需要保留的
sample_id列加入列表 - 遍历你的子串列表,找出所有列名包含对应子串的列,添加到列表中
- 最后用这个列列表直接从原DataFrame中筛选
import pandas as pd # 模拟你的原始数据 df1 = pd.DataFrame({ 'sample_id': [1], 'col9381.3': [2], 'col8371.8': [3], 'col71937.9': [4], 'col19993.1': [5] }) lst = ['col93','col71'] # 初始化要保留的列,先加入sample_id keep_cols = ['sample_id'] # 遍历子串,收集匹配的列 for substr in lst: matching_cols = [col for col in df1.columns if substr in col] keep_cols.extend(matching_cols) # 去重(避免同一列匹配多个子串的情况),同时保持原始列顺序 keep_cols = list(dict.fromkeys(keep_cols)) # 从原DataFrame筛选列 df_final = df1[keep_cols] print(df_final)
方法二:改进你的循环来合并结果
如果你想保留循环的思路,可以先初始化一个包含sample_id的结果DataFrame,然后每次把筛选出的列合并进去:
import pandas as pd df1 = pd.DataFrame({ 'sample_id': [1], 'col9381.3': [2], 'col8371.8': [3], 'col71937.9': [4], 'col19993.1': [5] }) lst = ['col93','col71'] # 初始化结果,先保留sample_id df_final = df1[['sample_id']].copy() for i in lst: df2 = df1.filter(regex=i) if df2.shape[1] > 0: # 按列合并筛选出的结果 df_final = pd.concat([df_final, df2], axis=1) print(df_final)
方法三:用正则表达式一次性匹配(最简洁)
可以把你的子串列表拼成一个正则表达式,用filter方法一次性筛选所有符合条件的列,再和sample_id合并:
import pandas as pd df1 = pd.DataFrame({ 'sample_id': [1], 'col9381.3': [2], 'col8371.8': [3], 'col71937.9': [4], 'col19993.1': [5] }) lst = ['col93','col71'] # 把子串拼成正则,|表示“或”的逻辑 regex_pattern = '|'.join(lst) # 合并sample_id和筛选出的列 df_final = pd.concat([df1[['sample_id']], df1.filter(regex=regex_pattern)], axis=1) print(df_final)
这三种方法都能得到你想要的结果,其中方法一和方法三更适合数据量大的场景,避免多次合并的开销。
内容的提问来源于stack exchange,提问作者John
相关产品推荐
相关产品推荐

