Python如何检测两个单词是否在指定词距内,可借助NLTK实现
问题解答
NLTK 现成函数说明
NLTK 没有直接实现该需求的专属封装函数,直接用原生 Python 实现即可,逻辑简单且运行效率足够应对常规分词结果的处理。
实现代码
先修正你给出的示例数据中的语法错误('uses,' 存在多余逗号和引号错位问题),完整可运行代码如下:
# 分词后的结果列表 file_contents = [ ['man', 'once', 'upon', 'time', 'love', 'princess'], ['python', 'code', 'cool', 'uses', 'java'], ['man', 'help', 'test', 'weird', 'love'] ] def check_words_within(word1: str, word2: str, size: int) -> list: result = [] for sub_list in file_contents: # 分别收集两个目标词在当前子列表中的所有索引 idx_list1 = [i for i, word in enumerate(sub_list) if word == word1] idx_list2 = [i for i, word in enumerate(sub_list) if word == word2] # 不存在任意一个目标词直接跳过 if not idx_list1 or not idx_list2: continue # 遍历所有索引对判断间隔是否符合要求 for i in idx_list1: for j in idx_list2: # 间隔单词数 = 两个位置的差绝对值 - 1 if abs(i - j) - 1 <= size: # 按示例要求返回两个词之间的切片(包含两个词本身) start = min(i, j) end = max(i, j) + 1 result.append(sub_list[start:end]) # 若一个子列表仅需返回第一个匹配结果,保留break;需要所有匹配结果可删除break break else: continue break return result
测试验证
运行你给出的调用示例,结果完全符合预期:
check_words_within('man','love', 4)返回[['man', 'once', 'upon', 'time', 'love'], ['man', 'help', 'test', 'weird', 'love']]check_words_within('man','upon', 1)返回[['man', 'once', 'upon']]check_words_within('man','document',4)返回[]
逻辑调整说明
如果你需要返回整个符合条件的子列表而非两个词之间的切片,只需将 result.append(sub_list[start:end]) 修改为 result.append(sub_list) 即可。
内容的提问来源于stack exchange,提问作者David R
相关产品推荐
相关产品推荐

