You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中对DataFrame列名执行子串提取操作?

从列名提取子串用于字符串匹配

嘿,我来帮你搞定这个问题!你想要从列名里动态提取子串(比如从comment1abc里拿到abc),代替硬编码写进str.contains里,其实很简单,分几种情况给你说:

1. 固定位置/长度的子串提取

如果你确定要提取的abc总是列名的最后3个字符,直接用字符串切片就可以——比固定索引[8:11]更灵活,就算前面的字符长度变了(比如列名是comment2abc或者comment123abc),[-3:]都能准确拿到最后3个字符:

# 先指定目标列名
target_column = 'comment1abc'
# 从列名提取后缀abc
match_substring = target_column[-3:]
# 用提取到的子串做匹配,注意:str.contains本身返回布尔值,不用加==True
result = df[target_column].str.contains(match_substring)

2. 动态匹配模式(非固定长度)

如果列名的格式是comment{N}abc(N是任意数字),或者你不确定子串的位置,想用正则精准提取,可以用re模块:

import re

target_column = 'comment1abc'
# 匹配列名末尾的abc(如果是其他后缀也可以改正则)
match = re.search(r'abc$', target_column)
if match:
    match_substring = match.group()
    result = df[target_column].str.contains(match_substring)

要是想提取数字后面的所有内容(比如列名是comment123xyz也能拿到xyz),可以把正则改成r'\d(\w+)$',然后取match.group(1)。

3. 批量处理多个符合格式的列

如果你的DataFrame里有多个类似commentXabc的列,想批量处理,可以先筛选出符合条件的列,再循环处理:

# 筛选所有以abc结尾的列
matching_columns = [col for col in df.columns if col.endswith('abc')]

for col in matching_columns:
    # 提取后缀
    match_substring = col[-3:]
    # 新增一列存储匹配结果
    df[f'{col}_has_substring'] = df[col].str.contains(match_substring)

最后提个小优化:df['col1'].str.contains('abc') == True其实可以简化成df['col1'].str.contains('abc'),因为str.contains返回的就是布尔类型的Series,直接用就够啦~

内容的提问来源于stack exchange,提问作者singularity2047

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:10:17