如何用Python/Pandas计算文本中重复单词的词间距?
计算文本中重复单词的词间距(Python/Pandas实现)
问题描述
需要获取文本中相同单词的词间距,例如单词“one”在目标文本中出现3次,首次与第二次出现间隔12个单词,需用Python或Pandas实现该间隔的计算。
解决方案
方法1:纯Python实现(无需Pandas)
通过正则分割文本为单词列表,定位目标单词的索引后计算相邻间隔:
import re long_string = "one are marked by the ()meta-characters. two They group together the expressions contained one inside them, and you can one repeat the contents of a group with a repeating qualifier, such as there" # 分割文本为纯单词列表(去除标点,匹配完整单词) words = re.findall(r'\b\w+\b', long_string) # 获取目标单词"one"所有出现的索引位置 target_indices = [idx for idx, word in enumerate(words) if word == 'one'] # 计算相邻出现的单词间隔数(索引差-1,排除目标单词本身) word_gaps = [target_indices[i+1] - target_indices[i] - 1 for i in range(len(target_indices)-1)] print("目标单词出现的索引:", target_indices) print("相邻单词间隔数:", word_gaps)
输出结果:
目标单词出现的索引: [0, 13, 17] 相邻单词间隔数: [12, 3]
方法2:Pandas实现
将逻辑封装为函数,通过apply方法处理Series对象:
import pandas as pd import re long_string = "one are marked by the ()meta-characters. two They group together the expressions contained one inside them, and you can one repeat the contents of a group with a repeating qualifier, such as there" my_text = pd.Series([long_string]) def get_word_gaps(text, target_word): # 分割为纯单词列表 words = re.findall(r'\b\w+\b', text) # 定位目标单词索引 indices = [idx for idx, word in enumerate(words) if word == target_word] # 计算间隔(不足2次出现则返回空列表) if len(indices) < 2: return [] return [indices[i+1] - indices[i] - 1 for i in range(len(indices)-1)] # 应用到Series result = my_text.apply(lambda x: get_word_gaps(x, 'one')) print(result)
输出结果:
0 [12, 3] dtype: object
说明
- 正则
\b\w+\b用于匹配完整单词,避免将包含目标词的其他词汇(如"someone")误判;若无需严格匹配单词边界,可改为re.split(r'\W+', long_string)分割文本。 - 词间距的核心逻辑:后一次出现的索引减去前一次索引,再减1(排除目标单词本身)。
内容的提问来源于stack exchange,提问作者user16858520
相关产品推荐
相关产品推荐

