如何用Matplotlib绘制特朗普演讲Top10高频词?代码问题求助
特朗普演讲高频词可视化问题解决
问题核心
直接将包含(词汇,频次)元组的列表lst传给plt.plot()时,Matplotlib无法识别这种复合结构,必须把词汇和频次拆分成独立序列才能正常绘图。另外你的文本合并逻辑存在错误,用df_1 + df_2是对DataFrame做元素级数值加法,而非拼接文本内容,会导致后续词频统计完全失真。
错误原因拆解
- 数据结构不匹配:
lst是类似[("said",757), ("great",693)]的元组列表,plt.plot()需要两个独立参数:x轴的类别(词汇)和y轴的数值(频次),直接传入复合列表会让Matplotlib无法解析。 - 文本合并逻辑错误:
+运算符对DataFrame执行的是元素级运算,不是拼接文本,会导致原始演讲内容被错误修改,词频统计结果完全错误。
正确实现步骤
1. 修复文本合并逻辑
用pd.concat()合并所有文本DataFrame,再将所有行的文本拼接成一个完整字符串:
# 用列表存储所有文件路径,循环读取避免重复代码 file_paths = [ "Trump_speeches/BattleCreekDec19_2019.txt", "Trump_speeches/BemidjiSep18_2020.txt", "Trump_speeches/CharlestonFeb28_2020.txt", "Trump_speeches/CharlotteMar2_2020.txt", "Trump_speeches/CincinnatiAug1_2019.txt", "Trump_speeches/ColoradorSpringsFeb20_2020.txt", "Trump_speeches/DallasOct17_2019.txt", "Trump_speeches/DesMoinesJan30_2020.txt", "Trump_speeches/FayettevilleSep9_2019.txt", "Trump_speeches/FayettevilleSep19_2020.txt", "Trump_speeches/FreelandSep10_2020.txt", "Trump_speeches/GreenvilleJul17_2019.txt", "Trump_speeches/HendersonSep13_2020.txt", "Trump_speeches/HersheyDec10_2019.txt" ] dfs = [pd.read_table(path) for path in file_paths] # 合并所有DataFrame df = pd.concat(dfs, ignore_index=True) # 拼接所有文本为单个字符串 full_text = ' '.join(df.iloc[:, 0].astype(str))
2. 从lst拆分x/y轴数据
用列表推导式从lst中分别提取词汇和频次:
x = [word for word, count in lst] y = [count for word, count in lst]
3. 选择合适图表并绘图
针对类别型的高频词数据,条形图比折线图更直观,绘图示例:
plt.figure(figsize=(10, 6)) plt.bar(x, y, color='#ff7f0e') # 添加图表细节,避免标签重叠 plt.title('特朗普集会演讲Top10高频词统计', fontsize=14) plt.xlabel('词汇', fontsize=12) plt.ylabel('出现频次', fontsize=12) plt.xticks(rotation=45, ha='right') plt.tight_layout() plt.show()
完整修正代码
import pandas as pd import operator import matplotlib.pyplot as plt # 读取并合并所有演讲文本 file_paths = [ "Trump_speeches/BattleCreekDec19_2019.txt", "Trump_speeches/BemidjiSep18_2020.txt", "Trump_speeches/CharlestonFeb28_2020.txt", "Trump_speeches/CharlotteMar2_2020.txt", "Trump_speeches/CincinnatiAug1_2019.txt", "Trump_speeches/ColoradorSpringsFeb20_2020.txt", "Trump_speeches/DallasOct17_2019.txt", "Trump_speeches/DesMoinesJan30_2020.txt", "Trump_speeches/FayettevilleSep9_2019.txt", "Trump_speeches/FayettevilleSep19_2020.txt", "Trump_speeches/FreelandSep10_2020.txt", "Trump_speeches/GreenvilleJul17_2019.txt", "Trump_speeches/HendersonSep13_2020.txt", "Trump_speeches/HersheyDec10_2019.txt" ] dfs = [pd.read_table(path) for path in file_paths] df = pd.concat(dfs, ignore_index=True) full_text = ' '.join(df.iloc[:, 0].astype(str)) # 完整停用词列表 def stopwords(): return ["","a","about","above","after","again","against","all","am","an","and","any","are","aren't","as","at","be","because","been","before","being","below","between","both","but","by","can't","cannot","could","couldn't","did","didn't","do","does","doesn't","doing","don't","down","during","each","few","for","from","further","had","hadn't","has","hasn't","have","haven't","having","he","he'd","he'll","he's","her","here","here's","hers","herself","him","himself","his","how","how's","i","i'd","i'll","i'm","i've","if","in","into","is","isn't","it","it's","its","itself","let's","me","more","most","mustn't","my","myself","no","nor","not","of","off","on","once","only","or","other","ought","our","ours","ourselves","out","over","own","same","shan't","she","she'd","she'll","she's","should","shouldn't","so","some","such","than","that","that's","the","their","theirs","them","themselves","then","there","there's","these","they","they'd","they'll","they're","they've","this","those","through","to","too","under","until","up","very","was","wasn't","we","we'd","we'll","we're","we've","were","weren't","what","what's","when","when's","where","where's","which","while","who","who's","whom","why","why's","with","won't","would","wouldn't","you","you'd","you'll","you're","you've","your","yours","yourself","yourselves"] def tokenizer(text, lower=True, stopword=True): if stopword: stopwordlist = stopwords() tokens = [token for token in text.lower().split(' ') if token not in stopwordlist] elif lower: tokens = text.lower().split() else: tokens = text.split() return tokens def word_count(text, sort = True, stopword=True): counter = dict() for word in tokenizer(text, stopword=stopword): counter.setdefault(word, 0) counter[word] += 1 if sort: counter = dict(sorted(counter.items(), key=operator.itemgetter(1), reverse=True)) return counter # 统计Top10高频词 wc = word_count(full_text) lst = list(wc.items())[:10] print('[INFO] 特朗普集会演讲Top10高频词:') print(lst) # 拆分x/y轴数据并绘图 x = [word for word, count in lst] y = [count for word, count in lst] plt.figure(figsize=(10, 6)) plt.bar(x, y, color='#ff7f0e') plt.title('特朗普集会演讲Top10高频词统计', fontsize=14) plt.xlabel('词汇', fontsize=12) plt.ylabel('出现频次', fontsize=12) plt.xticks(rotation=45, ha='right') plt.tight_layout() plt.show()
额外优化建议
- 用循环读取文件替代重复定义
df_1到df_14,代码更简洁易维护。 - 确保停用词列表完整(之前代码中的
.....需替换为完整停用词)。 - 若坚持用折线图,将
plt.bar()替换为plt.plot(x, y)即可,但条形图更适合类别数据对比。
内容的提问来源于stack exchange,提问作者Helten1997
相关产品推荐
相关产品推荐

