如何为Pandas DataFrame列的每行应用正则删除特定单词前的内容?
批量处理Pandas DataFrame文本列的正则替换方案
需求背景
你已经创建了如下重复内容的DataFrame:
import pandas as pd data = ['The text is interesting but short'] * 6 df = pd.DataFrame(data, columns=['Text'])
需要删除Text列每行中「interesting」之前的所有内容,且已实现单个字符串的正则处理逻辑,现在要将该逻辑批量应用到整列。
修正单个字符串处理逻辑
注意你原代码中up_to_word设为了"is",不符合需求,需改为目标单词"interesting",修正后的单字符串处理代码:
import re date_div = "The text is interesting but short" up_to_word = "interesting" rx_to_first = r'^.*?{}'.format(re.escape(up_to_word)) print(re.sub(rx_to_first, '', date_div, flags=re.DOTALL).strip()) # 输出:but short
批量应用到DataFrame列的两种方法
方法1:使用apply()逐行处理
将正则逻辑封装为函数,通过apply作用到Text列:
import pandas as pd import re # 创建原始DataFrame data = ['The text is interesting but short'] * 6 df = pd.DataFrame(data, columns=['Text']) def remove_before_target(text, target_word): rx = r'^.*?{}'.format(re.escape(target_word)) return re.sub(rx, '', text, flags=re.DOTALL).strip() # 生成处理后的新列 df['Processed_Text'] = df['Text'].apply(remove_before_target, target_word='interesting')
方法2:使用向量化str.replace()(推荐)
Pandas的字符串方法支持直接对整列执行正则替换,性能比apply更优,适合大数据量场景:
import pandas as pd import re data = ['The text is interesting but short'] * 6 df = pd.DataFrame(data, columns=['Text']) target_word = 'interesting' rx = r'^.*?{}'.format(re.escape(target_word)) # 直接处理整列并生成新列 df['Processed_Text'] = df['Text'].str.replace(rx, '', flags=re.DOTALL).str.strip()
内容的提问来源于stack exchange,提问作者user19783276
相关产品推荐
相关产品推荐

