Python pandas如何实现不区分大小写的字符串子串匹配并返回目标词
不区分大小写整词匹配实现方案
核心需求说明
- 匹配目标词时不区分大小写
- 仅匹配独立完整单词,避免词形变化的子串命中(例如
apples不能命中目标词apple)
实现代码
首先需要导入正则模块re用于边界匹配:
import pandas as pd import re # 原始数据 df_word_list = pd.DataFrame({'word': ['apple','cool']}) df = pd.DataFrame({'sentence': ['"Apple is cool"','"I like APPLE"','"I like apples"']}) # 预处理目标词表 origin_words = df_word_list['word'].tolist() lower_word_map = {w.lower(): w for w in origin_words} # 生成带单词边界的正则规则,开启忽略大小写模式 match_pattern = re.compile( r'\b(' + '|'.join(re.escape(w.lower()) for w in origin_words) + r')\b', flags=re.IGNORECASE ) # 定义匹配函数 def get_match_result(sentence): matched = match_pattern.findall(sentence) # 去重并映射回原词表的写法 res = list({lower_word_map[m.lower()] for m in matched}) return ', '.join(res) if res else '空' # 批量处理所有句子 df['result'] = df['sentence'].apply(get_match_result)
输出结果
执行后df['result']的输出完全符合预期:
- apple, cool
- apple
- 空
关键修改说明
- 用正则参数
re.IGNORECASE实现不区分大小写匹配,无需手动统一转换句子大小写 - 正则中加入
\b单词边界限定,仅匹配独立的完整单词,规避apples命中apple的问题 - 用字典映射保留原目标词的大小写格式,输出和词表存储的写法一致
内容的提问来源于stack exchange,提问作者inic72
相关产品推荐
相关产品推荐

