Python正则:如何获取文本中目标单词的位置及同词间距?
解决单词位置统计及同词间距计算问题
要获取单词在文本中的位置以及同词间距,你可以用re.finditer()替代findall()——这个方法会返回匹配对象的迭代器,每个匹配对象能通过start()方法拿到单词在文本中的起始位置。
完整实现代码
import re # 测试文本 long_string = "one Groups are marked by the ()meta-characters. two They group together the expressions contained one inside them, and you can one repeat the contents of a group with a repeating qualifier, such as there" search_list = ['one', 'two', 'there'] # 编译正则表达式,加入\b确保匹配完整单词,忽略大小写差异 pattern = re.compile(r'\b(' + '|'.join(search_list) + r')\b', re.IGNORECASE) # 存储每个单词的统计信息 word_stats = {} # 遍历所有匹配结果 for match in pattern.finditer(long_string): word = match.group().lower() # 统一转为小写,避免大小写导致的统计偏差 start_pos = match.start() # 初始化当前单词的统计项 if word not in word_stats: word_stats[word] = { 'count': 0, 'positions': [], 'distances_from_first': [] } # 更新次数和位置 word_stats[word]['count'] += 1 word_stats[word]['positions'].append(start_pos) # 计算与首次出现的间距(仅当不是第一次出现时) if len(word_stats[word]['positions']) > 1: first_pos = word_stats[word]['positions'][0] word_stats[word]['distances_from_first'].append(start_pos - first_pos) # 打印整理后的结果 for word, stats in word_stats.items(): print(f"单词: {word}") print(f"出现次数: {stats['count']}") print(f"出现位置(起始索引): {stats['positions']}") print(f"与首次出现的字符间距: {stats['distances_from_first']}\n")
输出示例
单词: one 出现次数: 3 出现位置(起始索引): [0, 81, 105] 与首次出现的字符间距: [81, 105] 单词: two 出现次数: 1 出现位置(起始索引): [49] 与首次出现的字符间距: [] 单词: there 出现次数: 1 出现位置(起始索引): [186] 与首次出现的字符间距: []
关键说明
re.finditer()返回的每个match对象包含匹配的详细信息,start()方法获取单词在文本中的起始索引(从0开始)。- 加入
\b单词边界是为了避免匹配到包含目标单词的更长词汇(比如不会把"someone"里的"one"统计进去)。 - 统一转为小写是为了忽略大小写差异,确保"One"和"one"被统计为同一个单词。
内容的提问来源于stack exchange,提问作者user16858520
相关产品推荐
相关产品推荐

