You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Matplotlib绘制特朗普演讲Top10高频词?代码问题求助

特朗普演讲高频词可视化问题解决

问题核心

直接将包含(词汇,频次)元组的列表lst传给plt.plot()时,Matplotlib无法识别这种复合结构,必须把词汇和频次拆分成独立序列才能正常绘图。另外你的文本合并逻辑存在错误,用df_1 + df_2是对DataFrame做元素级数值加法,而非拼接文本内容,会导致后续词频统计完全失真。

错误原因拆解

  1. 数据结构不匹配:lst是类似[("said",757), ("great",693)]的元组列表,plt.plot()需要两个独立参数:x轴的类别(词汇)和y轴的数值(频次),直接传入复合列表会让Matplotlib无法解析。
  2. 文本合并逻辑错误:+运算符对DataFrame执行的是元素级运算,不是拼接文本,会导致原始演讲内容被错误修改,词频统计结果完全错误。

正确实现步骤

1. 修复文本合并逻辑

用pd.concat()合并所有文本DataFrame,再将所有行的文本拼接成一个完整字符串:

# 用列表存储所有文件路径,循环读取避免重复代码
file_paths = [
    "Trump_speeches/BattleCreekDec19_2019.txt",
    "Trump_speeches/BemidjiSep18_2020.txt",
    "Trump_speeches/CharlestonFeb28_2020.txt",
    "Trump_speeches/CharlotteMar2_2020.txt",
    "Trump_speeches/CincinnatiAug1_2019.txt",
    "Trump_speeches/ColoradorSpringsFeb20_2020.txt",
    "Trump_speeches/DallasOct17_2019.txt",
    "Trump_speeches/DesMoinesJan30_2020.txt",
    "Trump_speeches/FayettevilleSep9_2019.txt",
    "Trump_speeches/FayettevilleSep19_2020.txt",
    "Trump_speeches/FreelandSep10_2020.txt",
    "Trump_speeches/GreenvilleJul17_2019.txt",
    "Trump_speeches/HendersonSep13_2020.txt",
    "Trump_speeches/HersheyDec10_2019.txt"
]
dfs = [pd.read_table(path) for path in file_paths]
# 合并所有DataFrame
df = pd.concat(dfs, ignore_index=True)
# 拼接所有文本为单个字符串
full_text = ' '.join(df.iloc[:, 0].astype(str))

2. 从lst拆分x/y轴数据

用列表推导式从lst中分别提取词汇和频次:

x = [word for word, count in lst]
y = [count for word, count in lst]

3. 选择合适图表并绘图

针对类别型的高频词数据,条形图比折线图更直观,绘图示例:

plt.figure(figsize=(10, 6))
plt.bar(x, y, color='#ff7f0e')
# 添加图表细节,避免标签重叠
plt.title('特朗普集会演讲Top10高频词统计', fontsize=14)
plt.xlabel('词汇', fontsize=12)
plt.ylabel('出现频次', fontsize=12)
plt.xticks(rotation=45, ha='right')
plt.tight_layout()
plt.show()

完整修正代码

import pandas as pd
import operator
import matplotlib.pyplot as plt

# 读取并合并所有演讲文本
file_paths = [
    "Trump_speeches/BattleCreekDec19_2019.txt",
    "Trump_speeches/BemidjiSep18_2020.txt",
    "Trump_speeches/CharlestonFeb28_2020.txt",
    "Trump_speeches/CharlotteMar2_2020.txt",
    "Trump_speeches/CincinnatiAug1_2019.txt",
    "Trump_speeches/ColoradorSpringsFeb20_2020.txt",
    "Trump_speeches/DallasOct17_2019.txt",
    "Trump_speeches/DesMoinesJan30_2020.txt",
    "Trump_speeches/FayettevilleSep9_2019.txt",
    "Trump_speeches/FayettevilleSep19_2020.txt",
    "Trump_speeches/FreelandSep10_2020.txt",
    "Trump_speeches/GreenvilleJul17_2019.txt",
    "Trump_speeches/HendersonSep13_2020.txt",
    "Trump_speeches/HersheyDec10_2019.txt"
]
dfs = [pd.read_table(path) for path in file_paths]
df = pd.concat(dfs, ignore_index=True)
full_text = ' '.join(df.iloc[:, 0].astype(str))

# 完整停用词列表
def stopwords():
    return ["","a","about","above","after","again","against","all","am","an","and","any","are","aren't","as","at","be","because","been","before","being","below","between","both","but","by","can't","cannot","could","couldn't","did","didn't","do","does","doesn't","doing","don't","down","during","each","few","for","from","further","had","hadn't","has","hasn't","have","haven't","having","he","he'd","he'll","he's","her","here","here's","hers","herself","him","himself","his","how","how's","i","i'd","i'll","i'm","i've","if","in","into","is","isn't","it","it's","its","itself","let's","me","more","most","mustn't","my","myself","no","nor","not","of","off","on","once","only","or","other","ought","our","ours","ourselves","out","over","own","same","shan't","she","she'd","she'll","she's","should","shouldn't","so","some","such","than","that","that's","the","their","theirs","them","themselves","then","there","there's","these","they","they'd","they'll","they're","they've","this","those","through","to","too","under","until","up","very","was","wasn't","we","we'd","we'll","we're","we've","were","weren't","what","what's","when","when's","where","where's","which","while","who","who's","whom","why","why's","with","won't","would","wouldn't","you","you'd","you'll","you're","you've","your","yours","yourself","yourselves"]

def tokenizer(text, lower=True, stopword=True):
    if stopword:
        stopwordlist = stopwords()
        tokens = [token for token in text.lower().split(' ') if token not in stopwordlist]
    elif lower:
        tokens = text.lower().split()
    else:
        tokens = text.split()
    return tokens

def word_count(text, sort = True, stopword=True):
    counter = dict()
    for word in tokenizer(text, stopword=stopword):
        counter.setdefault(word, 0)
        counter[word] += 1
    if sort:
        counter = dict(sorted(counter.items(), key=operator.itemgetter(1), reverse=True))
    return counter

# 统计Top10高频词
wc = word_count(full_text)
lst = list(wc.items())[:10]
print('[INFO] 特朗普集会演讲Top10高频词:')
print(lst)

# 拆分x/y轴数据并绘图
x = [word for word, count in lst]
y = [count for word, count in lst]

plt.figure(figsize=(10, 6))
plt.bar(x, y, color='#ff7f0e')
plt.title('特朗普集会演讲Top10高频词统计', fontsize=14)
plt.xlabel('词汇', fontsize=12)
plt.ylabel('出现频次', fontsize=12)
plt.xticks(rotation=45, ha='right')
plt.tight_layout()
plt.show()

额外优化建议

  • 用循环读取文件替代重复定义df_1到df_14,代码更简洁易维护。
  • 确保停用词列表完整(之前代码中的.....需替换为完整停用词)。
  • 若坚持用折线图,将plt.bar()替换为plt.plot(x, y)即可,但条形图更适合类别数据对比。

内容的提问来源于stack exchange,提问作者Helten1997

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 00:05:17