You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python找出加密货币推特数据中的高频词汇?

Hey there! 既然你已经搞定了加密货币推文的数据清洗和情感分析,现在要挖最常用的词汇对吧?我给你整理了几个实用的方案,适配你的场景:

提取加密货币推文高频词汇方案

1. 先把基础库补全

你开头的导入代码可以完善成这样,加上处理词频和文本需要的工具:

import pandas as pd
from collections import Counter
import re
# 处理英文停用词需要nltk库
import nltk
from nltk.corpus import stopwords
# 第一次用需要下载停用词数据集
nltk.download('stopwords')

2. 加载清洗后的CSV数据

假设你的CSV里有一列叫tweet_text(存的是已经清洗好的推文内容),先把数据加载进来:

# 读取清洗后的CSV文件
df = pd.read_csv('your_cleaned_tweets.csv')
# 提取推文文本列,顺便去掉空值
clean_tweets = df['tweet_text'].dropna().tolist()

3. 三种统计高频词汇的方法

方法一:用Counter快速统计(基础版)

适合快速得到结果,先把所有推文合并,分词后过滤噪音再统计:

# 把所有推文合并成一个大文本
full_text = ' '.join(clean_tweets)
# 分词并转小写(避免大小写导致重复统计,比如Bitcoin和bitcoin算同一个词)
words = re.findall(r'\b\w+\b', full_text.lower())
# 加载英文停用词,过滤掉无意义的词
stop_words = set(stopwords.words('english'))
# 过滤停用词和长度小于3的短词(比如a、the这类)
filtered_words = [word for word in words if word not in stop_words and len(word) > 2]
# 统计词频
word_freq = Counter(filtered_words)
# 获取前20个高频词
top_20_words = word_freq.most_common(20)

# 打印结果
print("Top 20 most frequent words:")
for word, count in top_20_words:
    print(f"{word}: {count}")

方法二:适配加密货币领域的优化版

因为你的数据是加密货币相关的,可能需要自定义过滤规则,比如排除转发标记rt、amp这类无意义词,同时保留btc、eth这类领域专属缩写:

# 自定义停用词集合,把通用停用词和领域无关词合并
custom_stopwords = stop_words.union({'rt', 'amp', 'crypto', 'cryptocurrency'})
# 重新过滤词汇
filtered_words = [word for word in words if word not in custom_stopwords and len(word) > 2]
# 重新统计
word_freq = Counter(filtered_words)
top_20_words = word_freq.most_common(20)

方法三:用词云可视化(可选)

如果想更直观展示高频词,可以生成词云图:

from wordcloud import WordCloud
import matplotlib.pyplot as plt

# 生成词云
wordcloud = WordCloud(
    width=800, 
    height=400, 
    background_color='white',
    max_words=50  # 最多展示50个词
).generate_from_frequencies(word_freq)

# 展示词云
plt.figure(figsize=(10, 5))
plt.imshow(wordcloud, interpolation='bilinear')
plt.axis('off')  # 隐藏坐标轴
plt.show()

几个小提醒

  • 确保你的数据已经完成去重、删除链接、删除@用户、去除特殊符号这些清洗步骤,不然统计结果会有很多噪音
  • 如果推文里有加密货币的专属词汇或缩写,别误当成停用词过滤掉,可以提前整理一个领域词列表做保留
  • 可以调整most_common()里的数字,获取更多或更少的高频词

内容的提问来源于stack exchange,提问作者Aziz Bokhari

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:54:08