You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为Pandas DataFrame列的每行应用正则删除特定单词前的内容?

批量处理Pandas DataFrame文本列的正则替换方案

需求背景

你已经创建了如下重复内容的DataFrame:

import pandas as pd

data = ['The text is interesting but short'] * 6
df = pd.DataFrame(data, columns=['Text'])

需要删除Text列每行中「interesting」之前的所有内容,且已实现单个字符串的正则处理逻辑,现在要将该逻辑批量应用到整列。

修正单个字符串处理逻辑

注意你原代码中up_to_word设为了"is",不符合需求,需改为目标单词"interesting",修正后的单字符串处理代码:

import re

date_div = "The text is interesting but short"
up_to_word = "interesting"
rx_to_first = r'^.*?{}'.format(re.escape(up_to_word))
print(re.sub(rx_to_first, '', date_div, flags=re.DOTALL).strip())
# 输出:but short

批量应用到DataFrame列的两种方法

方法1:使用apply()逐行处理

将正则逻辑封装为函数,通过apply作用到Text列:

import pandas as pd
import re

# 创建原始DataFrame
data = ['The text is interesting but short'] * 6
df = pd.DataFrame(data, columns=['Text'])

def remove_before_target(text, target_word):
    rx = r'^.*?{}'.format(re.escape(target_word))
    return re.sub(rx, '', text, flags=re.DOTALL).strip()

# 生成处理后的新列
df['Processed_Text'] = df['Text'].apply(remove_before_target, target_word='interesting')

方法2:使用向量化str.replace()(推荐)

Pandas的字符串方法支持直接对整列执行正则替换,性能比apply更优,适合大数据量场景:

import pandas as pd
import re

data = ['The text is interesting but short'] * 6
df = pd.DataFrame(data, columns=['Text'])

target_word = 'interesting'
rx = r'^.*?{}'.format(re.escape(target_word))
# 直接处理整列并生成新列
df['Processed_Text'] = df['Text'].str.replace(rx, '', flags=re.DOTALL).str.strip()

内容的提问来源于stack exchange,提问作者user19783276

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 13:00:56