You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python的NLTK获取含指定单词的句子索引?

解决方法

错误原因

  • list.index(x) 方法要求参数x是列表里的单个元素,但你传入的matches是一堆句子组成的列表,直接调用肯定报错。
  • 原代码还有冗余问题:如果一个句子同时包含多个搜索词(比如最后一句既有"Ipsum"又有"Aldus"),会被重复添加到matches里,结果变啰嗦。

修正后的代码

直接在遍历句子时同步记录索引和内容,一步到位解决问题:

from nltk.tokenize import sent_tokenize

text = "Lorem Ipsum is simply dummy text of the printing and typesetting industry. Lorem Ipsum has been the industry's standard dummy text ever since the 1500s, when an unknown printer took a galley of type and scrambled it to make a type specimen book. It has survived not only five centuries, but also the leap into electronic typesetting, remaining essentially unchanged. It was popularised in the 1960s with the release of Letraset sheets containing Lorem Ipsum passages, and more recently with desktop publishing software like Aldus PageMaker including versions of Lorem Ipsum."
search_words = ["Ipsum", "Aldus"]
matches = []
sentences = sent_tokenize(text)

# 遍历句子时同时拿到索引和内容
for idx, sentence in enumerate(sentences):
    # 检查句子是否含任意一个搜索词
    if any(word in sentence for word in search_words):
        # 存索引(要从1开始计数就写idx+1)和句子
        matches.append((idx, sentence))

# 输出结果
for idx, sent in matches:
    print(f"索引 {idx}: {sent}")

代码说明

  • enumerate()是Python里拿元素索引的标准操作,遍历列表时能同时拿到位置序号和对应内容。
  • any(...)用来判断句子是否包含目标词,避免同一句子被重复添加。
  • 要是习惯索引从1开始(比如日常说第1句、第2句),把代码里的idx改成idx+1就行。

原思路的问题

如果非要先收集句子再找索引,得注意:

  1. 要逐个遍历matches里的句子,单独调用.index(),但这个方法只会返回第一个匹配的索引,要是有重复句子会出错。
  2. 得先给matches去重,不然会重复处理。但这种方法既麻烦又容易出问题,不如直接用enumerate()高效。

内容的提问来源于stack exchange,提问作者user9299756

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 15:15:38