You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现正则匹配DataFrame文本时兼容element的缩写字典?

正则匹配DataFrame文本中的元素及缩写需求解决

需求说明

现有包含自由文本字段text的DataFrame,需通过正则表达式识别指定element元素。部分元素存在缩写形式,已生成缩写字典,要求修改代码实现:当元素在缩写字典中时,正则同时匹配该元素的全称及其所有缩写(例如缩写ca匹配全称cat)。


现有数据示例

customerId                text element  code
0           1  Something with Cat     cat     0
1           3  That is a huge dog     dog     1
2           3         Hello agian   mouse     2
3           3        This is a ca     cat     0

当前代码

import pandas as pd
import re

d = {
    "customerId": [1, 3, 3, 3],
    "text": ["Something with Cat", "That is a huge dog", "Hello agian", 'This is a ca'],
    "element": ['cat', 'dog', 'mouse', 'cat'],
    "code": [9,8,7, 9]
}
df = pd.DataFrame(data=d)
df['code'] = df['element'].astype('category').cat.codes
print(df)

abbreviation = {
    "cat": {
        "abbrev1": "ca",
    },
} 

%%time

elements = df['element'].unique()
def f(x):
    match = 999
    for element in elements:
        elements2 = [element]
        y = bool(re.search(element, x['text'], re.IGNORECASE))
        if(y):
            match = x['code']
            break
    x['test'] = match
    return x
df['test'] = None
df = df.apply(lambda x: f(x), axis = 1)

当前运行结果

customerId                text element  code  test
0           1  Something with Cat     cat     0     0
1           3  That is a huge dog     dog     1     1
2           3         Hello agian   mouse     2   999
3           3        This is a ca     cat     0   999

期望结果

customerId                text element  code  test
0           1  Something with Cat     cat     0     0
1           3  That is a huge dog     dog     1     1
2           3         Hello agian   mouse     2   999
3           3        This is a ca     cat     0     0

修改后的代码及说明

核心优化点

  1. 预构建每个元素的正则匹配模式,包含全称+所有缩写
  2. 使用re.escape()处理元素文本,避免特殊字符干扰正则语法
  3. 精准匹配当前行对应的元素模式,提升逻辑效率

完整修改代码

import pandas as pd
import re

d = {
    "customerId": [1, 3, 3, 3],
    "text": ["Something with Cat", "That is a huge dog", "Hello agian", 'This is a ca'],
    "element": ['cat', 'dog', 'mouse', 'cat'],
    "code": [9,8,7, 9]
}
df = pd.DataFrame(data=d)
df['code'] = df['element'].astype('category').cat.codes
print(df)

abbreviation = {
    "cat": {
        "abbrev1": "ca",
    },
} 

%%time

# 预生成每个元素的正则匹配模式(全称+缩写)
element_patterns = {}
for elem in df['element'].unique():
    # 基础匹配项:元素本身
    patterns = [re.escape(elem)]
    # 添加所有缩写项(如果存在)
    if elem in abbreviation:
        patterns.extend(re.escape(abbrev) for abbrev in abbreviation[elem].values())
    # 构建忽略大小写的正则模式
    element_patterns[elem] = re.compile('|'.join(patterns), re.IGNORECASE)

def f(x):
    match = 999
    target_elem = x['element']
    # 获取当前元素对应的匹配模式
    pattern = element_patterns.get(target_elem)
    if pattern and pattern.search(x['text']):
        match = x['code']
    x['test'] = match
    return x

df['test'] = None
df = df.apply(f, axis=1)
print(df)

代码解释

  • 预生成正则模式避免循环内重复构建,提升运行效率
  • re.escape()确保元素中的特殊字符(如., *)不会被解析为正则语法
  • 仅匹配当前行对应的元素模式,逻辑更精准,避免无关元素的误匹配
  • 保留忽略大小写的匹配规则,兼容不同大小写的文本内容

内容的提问来源于stack exchange,提问作者Test

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 01:15:38