You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python/Pandas计算文本中重复单词的词间距?

计算文本中重复单词的词间距(Python/Pandas实现)

问题描述

需要获取文本中相同单词的词间距,例如单词“one”在目标文本中出现3次,首次与第二次出现间隔12个单词,需用Python或Pandas实现该间隔的计算。

解决方案

方法1:纯Python实现(无需Pandas)

通过正则分割文本为单词列表,定位目标单词的索引后计算相邻间隔:

import re

long_string = "one are marked by the ()meta-characters. two They group together the expressions contained one inside them, and you can one repeat the contents of a group with a repeating qualifier, such as there"

# 分割文本为纯单词列表(去除标点,匹配完整单词)
words = re.findall(r'\b\w+\b', long_string)
# 获取目标单词"one"所有出现的索引位置
target_indices = [idx for idx, word in enumerate(words) if word == 'one']
# 计算相邻出现的单词间隔数(索引差-1,排除目标单词本身)
word_gaps = [target_indices[i+1] - target_indices[i] - 1 for i in range(len(target_indices)-1)]

print("目标单词出现的索引:", target_indices)
print("相邻单词间隔数:", word_gaps)

输出结果:

目标单词出现的索引: [0, 13, 17]
相邻单词间隔数: [12, 3]

方法2:Pandas实现

将逻辑封装为函数,通过apply方法处理Series对象:

import pandas as pd
import re

long_string = "one are marked by the ()meta-characters. two They group together the expressions contained one inside them, and you can one repeat the contents of a group with a repeating qualifier, such as there"
my_text = pd.Series([long_string])

def get_word_gaps(text, target_word):
    # 分割为纯单词列表
    words = re.findall(r'\b\w+\b', text)
    # 定位目标单词索引
    indices = [idx for idx, word in enumerate(words) if word == target_word]
    # 计算间隔(不足2次出现则返回空列表)
    if len(indices) < 2:
        return []
    return [indices[i+1] - indices[i] - 1 for i in range(len(indices)-1)]

# 应用到Series
result = my_text.apply(lambda x: get_word_gaps(x, 'one'))
print(result)

输出结果:

0    [12, 3]
dtype: object

说明

  • 正则\b\w+\b用于匹配完整单词,避免将包含目标词的其他词汇(如"someone")误判;若无需严格匹配单词边界,可改为re.split(r'\W+', long_string)分割文本。
  • 词间距的核心逻辑:后一次出现的索引减去前一次索引,再减1(排除目标单词本身)。

内容的提问来源于stack exchange,提问作者user16858520

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 07:24:22