You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python pandas如何实现不区分大小写的字符串子串匹配并返回目标词

不区分大小写整词匹配实现方案

核心需求说明

  • 匹配目标词时不区分大小写
  • 仅匹配独立完整单词,避免词形变化的子串命中(例如apples不能命中目标词apple)

实现代码

首先需要导入正则模块re用于边界匹配:

import pandas as pd
import re

# 原始数据
df_word_list = pd.DataFrame({'word':  ['apple','cool']})
df = pd.DataFrame({'sentence':  ['"Apple is cool"','"I like APPLE"','"I like apples"']})

# 预处理目标词表
origin_words = df_word_list['word'].tolist()
lower_word_map = {w.lower(): w for w in origin_words}
# 生成带单词边界的正则规则,开启忽略大小写模式
match_pattern = re.compile(
    r'\b(' + '|'.join(re.escape(w.lower()) for w in origin_words) + r')\b',
    flags=re.IGNORECASE
)

# 定义匹配函数
def get_match_result(sentence):
    matched = match_pattern.findall(sentence)
    # 去重并映射回原词表的写法
    res = list({lower_word_map[m.lower()] for m in matched})
    return ', '.join(res) if res else '空'

# 批量处理所有句子
df['result'] = df['sentence'].apply(get_match_result)

输出结果

执行后df['result']的输出完全符合预期:

  1. apple, cool
  2. apple
  3. 空

关键修改说明

  • 用正则参数re.IGNORECASE实现不区分大小写匹配,无需手动统一转换句子大小写
  • 正则中加入\b单词边界限定,仅匹配独立的完整单词,规避apples命中apple的问题
  • 用字典映射保留原目标词的大小写格式,输出和词表存储的写法一致

内容的提问来源于stack exchange,提问作者inic72

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 23:45:03