You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在无空格句子中查找目标词的精确索引(支持重复匹配)

无空格句子匹配目标单词精确索引的解决方案

实现思路

利用re.finditer()结合正则表达式的精确匹配逻辑,遍历每一行的句子与目标单词,收集所有匹配的起始索引,完全规避split方法的使用。

完整代码

import pandas as pd
import re

# 初始化示例DataFrame
df = pd.DataFrame({
    'sentence': {0: 'idontlikeanapple.dislike', 1: 'welcometomyroom.plzcomein', 2: 'thisisthetable'},
    'target': {0: 'like', 1: 'come', 2: 'is'}
})

def get_target_indices(row):
    sentence = row['sentence']
    target = row['target']
    # 转义目标单词中的特殊字符,避免正则语法冲突
    regex_pattern = re.compile(re.escape(target))
    # 遍历所有匹配项,提取起始索引
    return [match.start() for match in regex_pattern.finditer(sentence)]

# 给DataFrame新增索引列
df['match_indices'] = df.apply(get_target_indices, axis=1)

print(df)

关键细节说明

  • re.escape(target):处理目标单词中的特殊字符(比如示例里的.),防止其被当作正则语法解析。
  • finditer():返回所有非重叠匹配的迭代器,每个匹配对象的start()方法直接给出匹配的起始位置。
  • apply(axis=1):针对DataFrame每一行执行匹配逻辑,保证每个句子对应自身的目标单词。

运行输出

sentence target match_indices
0  idontlikeanapple.dislike    like      [5, 20]
1  welcometomyroom.plzcomein   come      [15, 19]
2           thisisthetable     is          [2]

内容的提问来源于stack exchange,提问作者jungmin heo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 18:02:41