如何匹配字符串中的目标短语并获取对应单词索引?
获取匹配短语的单词索引解决方案
嘿,我懂你的痛点——你要的是单词层面的索引,而不是字符位置,用re.search确实解决不了这个问题。咱们换个思路,直接从拆分后的单词列表入手就简单多了,下面是具体的实现方法:
基础实现(精确匹配)
先看最直接的情况,假设你需要完全匹配单词(包括标点):
original_str = "This is an example sentence, it is for demonstration only" target_phrase = "example sentence" # 拆分原字符串和目标短语为单词列表 splitted_words = original_str.split() target_words = target_phrase.split() target_len = len(target_words) match_indices = None # 遍历原单词列表,寻找连续匹配的子序列 for i in range(len(splitted_words) - target_len + 1): if splitted_words[i:i+target_len] == target_words: start_idx = i end_idx = i + target_len - 1 match_indices = (start_idx, end_idx) break if match_indices: print(f"找到匹配!起始单词索引:{match_indices[0]},结束单词索引:{match_indices[1]}") else: print("未找到匹配的短语")
不过注意哦,你例子里原字符串的sentence,带逗号,而目标短语里是sentence,这会导致上面的代码匹配失败。如果需要忽略标点或者大小写,可以看下面的优化版本。
优化版(忽略标点/大小写)
如果要处理标点、大小写不一致的情况,我们可以先对每个单词做清洗:
import string def clean_word(word): # 去除单词两端的标点,同时转成小写(如果需要大小写不敏感匹配) return word.strip(string.punctuation).lower() original_str = "This is an example sentence, it is for demonstration only" target_phrase = "example sentence" splitted_words = original_str.split() target_words = target_phrase.split() # 清洗所有单词 cleaned_original = [clean_word(word) for word in splitted_words] cleaned_target = [clean_word(word) for word in target_words] target_len = len(cleaned_target) match_indices = None for i in range(len(cleaned_original) - target_len + 1): if cleaned_original[i:i+target_len] == cleaned_target: start_idx = i end_idx = i + target_len - 1 match_indices = (start_idx, end_idx) break if match_indices: print(f"找到匹配!起始单词索引:{match_indices[0]},结束单词索引:{match_indices[1]}") else: print("未找到匹配的短语")
这个版本就能正确匹配到你例子里的example(索引3)和sentence,(索引4)了。
为什么不用正则?
re.search返回的是字符的起始位置,要转成单词索引的话,你得统计该位置之前有多少个空格,这种方法不仅麻烦,还容易在单词间有多个空格的情况下出错。直接操作拆分后的单词列表,逻辑更清晰,也更可靠。
内容的提问来源于stack exchange,提问作者Avinash
相关产品推荐
相关产品推荐

