如何在DataFrame文本列中匹配Series子串并新增结果列?
解决方法
要实现从DataFrame的text列中提取匹配指定子串的内容,可以通过正则表达式捕获组结合str.extract()方法完成,具体步骤如下:
1. 准备示例数据
import pandas as pd # 构建目标DataFrame df = pd.DataFrame({ 'text': ['abdcdtext1wrew', 'qwerqdtext2cvufu', 'iuotext3tvbv', 'iuotvbvewre'], 'sth': ['...', '...', '...', '...'] }) # 构建存储匹配子串的Series df_look_for = pd.Series(['text1', 'text2', 'text3'], name='look_for')
2. 生成匹配正则模式
将需要匹配的子串用|拼接,包裹在正则捕获组()中,让str.extract()可以提取到具体匹配的内容:
# 生成正则匹配模式 pattern = '(' + '|'.join(df_look_for) + ')'
3. 提取匹配子串并添加新列
使用str.extract()提取匹配内容,无匹配项会自动填充NaN:
df['found_str'] = df['text'].str.extract(pattern, expand=False)
最终结果
执行后你的DataFrame会变成:
| text | sth | found_str |
|---|---|---|
| abdcdtext1wrew | ... | text1 |
| qwerqdtext2cvufu | ... | text2 |
| iuotext3tvbv | ... | text3 |
| iuotvbvewre | ... | NaN |
补充说明
你之前用str.contains()未成功的原因是:该方法仅返回布尔值(是否包含子串),无法提取具体匹配的子串。而str.extract()通过正则捕获组,可以直接获取到匹配的内容,完全满足你的需求。
如果存在单个text匹配多个子串的场景,上述方法会返回第一个匹配的子串;若需要提取所有匹配项,可以改用str.findall()后再做格式处理。
内容的提问来源于stack exchange,提问作者complog
相关产品推荐
相关产品推荐

