You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在自定义文本数据上调用NLTK的collocation_list()识别搭配

NLTK自有数据调用collocation_list()报错解决方案

报错根因

你遇到的AttributeError: 'list' object has no attribute 'collocation_list'报错属于对象类型不匹配问题:

NLTK示例中的text4是nltk.text.Text类的内置实例,该类原生封装了collocation_list()搭配识别方法;而你通过word_tokenize()分词得到的Tokens是普通Python列表,列表类本身没有该方法,因此调用报错。

解决方法

只需额外一步操作:将分词得到的列表包装为nltk.text.Text类的实例,再调用对应方法即可。

修正后完整可运行代码

import nltk
from nltk import word_tokenize
from nltk.text import Text
from nltk.probability import FreqDist

# 读取文件
File1 = open("/Applications/Python 3.9/StormZuluStory.txt",encoding="Latin-1")
StormZuluStory=File1.read()
File1.close() # 补充关闭文件操作,避免资源泄漏

File2 = open("/Applications/Python 3.9/StormZuluPOSStory.txt",encoding="Latin-1")
StormZuluPOSStory=File2.read()
File2.close()

# 分词并包装为NLTK Text实例
Tokens = word_tokenize(StormZuluStory)
text_obj = Text(Tokens) # 核心修改:将列表转为Text类实例

# 原有频次统计逻辑
fdist = FreqDist(Tokens)
Freq1 = fdist.most_common(30)
print(Freq1)
Plot1 = fdist.plot(30,cumulative=True)

# 调用搭配识别方法
collocations = text_obj.collocation_list()
print(collocations)

可选参数说明

collocation_list()支持自定义参数调整识别规则,常用参数如下:

  • num: 返回的搭配数量,默认20
  • window_size: 搭配统计的窗口大小,默认2
  • scoring: 搭配评分算法,默认使用似然比likelihood_ratio,也可选择pmi等算法

内容的提问来源于stack exchange,提问作者Lindsay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 18:39:03