You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用matplotlib可视化NLTK提取的高频bigram频次数据

报错原因

negbigram.plot.barh()触发AttributeError: 'list' object has no attribute 'plot'的核心原因:negbigram是Python原生列表对象,不支持pandas数据结构的内置绘图方法,也不是可直接传入绘图接口的结构化数组,因此无法直接调用plot类属性。

实现方案

先把列表里存储的「词对+频次」结构拆分成绘图可用的标签、数值两部分,即可用matplotlib完成可视化,以下提供两种可直接运行的实现方式:

  • 方式1:纯matplotlib原生实现
    无需额外依赖pandas,直接拆分数据绘图即可:

import matplotlib.pyplot as plt

拆分词对标签、频次数值:把元组形式的词对拼接为空格分隔的字符串

bigram_text = [' '.join(word_pair) for word_pair, count in negbigram]
count_data = [count for word_pair, count in negbigram]

反转列表顺序,让频次最高的词对显示在条形图最顶部

bigram_text.reverse()
count_data.reverse()

绘制水平条形图

plt.figure(figsize=(12, 8))
plt.barh(bigram_text, count_data, color='steelblue', height=0.9)
plt.xlabel('出现频次')
plt.ylabel('高频Bigram词对')
plt.title('负面评论Top5高频二元词组统计')

在每个条形末端标注具体频次数值

for idx, val in enumerate(count_data):
plt.text(val + 8, idx, str(val), va='center')

plt.tight_layout()
plt.show()

- 方式2:转pandas结构使用原有绘图语法
如果需要沿用`.plot.barh()`的写法,只需要把原生列表转成pandas的Series结构即可:
```python
import pandas as pd
import matplotlib.pyplot as plt

# 构造pandas Series,索引为词对文本,值为对应频次
negbigram_s = pd.Series(
  data=[count for word_pair, count in negbigram],
  index=[' '.join(word_pair) for word_pair, count in negbigram]
).sort_values()  # 升序排列保证最高频项在图的顶部

# 直接调用原有绘图逻辑
negbigram_s.plot.barh(color='steelblue', width=0.9, figsize=(12,8))
plt.xlabel('出现频次')
plt.title('负面评论Top5高频二元词组统计')
plt.tight_layout()
plt.show()

如果需要生成统计表格,直接把结果转成pandas DataFrame即可输出规整的两列结果:

import pandas as pd
bigram_df = pd.DataFrame(negbigram, columns=['Bigram词对', '出现频次'])
bigram_df['Bigram词对'] = bigram_df['Bigram词对'].apply(lambda x: ' '.join(x))
print(bigram_df)

运行后输出的表格格式如下:

Bigram词对出现频次
0stopped working637
1battery life408
2waste money354
3samsung galaxy322
4apple store289

内容的提问来源于stack exchange,提问作者kingcameron

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 00:54:32