如何使用matplotlib可视化NLTK提取的高频bigram频次数据
negbigram.plot.barh()触发AttributeError: 'list' object has no attribute 'plot'的核心原因:negbigram是Python原生列表对象,不支持pandas数据结构的内置绘图方法,也不是可直接传入绘图接口的结构化数组,因此无法直接调用plot类属性。
先把列表里存储的「词对+频次」结构拆分成绘图可用的标签、数值两部分,即可用matplotlib完成可视化,以下提供两种可直接运行的实现方式:
- 方式1:纯matplotlib原生实现
无需额外依赖pandas,直接拆分数据绘图即可:
import matplotlib.pyplot as plt
bigram_text = [' '.join(word_pair) for word_pair, count in negbigram]
count_data = [count for word_pair, count in negbigram]
bigram_text.reverse()
count_data.reverse()
plt.figure(figsize=(12, 8))
plt.barh(bigram_text, count_data, color='steelblue', height=0.9)
plt.xlabel('出现频次')
plt.ylabel('高频Bigram词对')
plt.title('负面评论Top5高频二元词组统计')
for idx, val in enumerate(count_data):
plt.text(val + 8, idx, str(val), va='center')
plt.tight_layout()
plt.show()
- 方式2:转pandas结构使用原有绘图语法 如果需要沿用`.plot.barh()`的写法,只需要把原生列表转成pandas的Series结构即可: ```python import pandas as pd import matplotlib.pyplot as plt # 构造pandas Series,索引为词对文本,值为对应频次 negbigram_s = pd.Series( data=[count for word_pair, count in negbigram], index=[' '.join(word_pair) for word_pair, count in negbigram] ).sort_values() # 升序排列保证最高频项在图的顶部 # 直接调用原有绘图逻辑 negbigram_s.plot.barh(color='steelblue', width=0.9, figsize=(12,8)) plt.xlabel('出现频次') plt.title('负面评论Top5高频二元词组统计') plt.tight_layout() plt.show()
如果需要生成统计表格,直接把结果转成pandas DataFrame即可输出规整的两列结果:
import pandas as pd bigram_df = pd.DataFrame(negbigram, columns=['Bigram词对', '出现频次']) bigram_df['Bigram词对'] = bigram_df['Bigram词对'].apply(lambda x: ' '.join(x)) print(bigram_df)
运行后输出的表格格式如下:
Bigram词对 出现频次 0 stopped working 637 1 battery life 408 2 waste money 354 3 samsung galaxy 322 4 apple store 289
内容的提问来源于stack exchange,提问作者kingcameron

