如何基于TextBlob提取Facebook帖子数据中的高频名词、形容词及名词短语?
嘿,这事儿好办!你已经用TextBlob完成了基础的词性标注和名词短语提取,接下来要统计高频的名词、形容词和名词短语,其实用Python的collections.Counter就能轻松搞定,给你一步步来:
1. 提取最常见的名词和形容词
TextBlob返回的tags是(单词, 词性标签)的元组列表,我们需要先筛选出对应词性的词汇,再统计词频:
首先导入必要的工具:
from collections import Counter
然后处理词性筛选和统计:
# 定义需要筛选的词性标签(参考Penn Treebank词性标注体系) noun_tags = {'NN', 'NNS', 'NNP', 'NNPS'} # 单数/复数普通名词、专有名词 adj_tags = {'JJ', 'JJR', 'JJS'} # 原级/比较级/最高级形容词 # 把所有帖子的词性标注展开成一个大列表 all_tags = [] for post_tags in data_tags: all_tags.extend(post_tags) # 提取名词并统一小写(避免大小写差异导致统计偏差) nouns = [word.lower() for word, tag in all_tags if tag in noun_tags] # 提取形容词并统一小写 adjectives = [word.lower() for word, tag in all_tags if tag in adj_tags] # 统计前20个最常见的名词和形容词 top_nouns = Counter(nouns).most_common(20) top_adjectives = Counter(adjectives).most_common(20) # 输出结果 print("Top 20 常见名词:", top_nouns) print("Top 20 常见形容词:", top_adjectives)
2. 提取最常见的名词短语
对于已经提取好的data_noun_phrases,我们只需要把所有帖子的短语合并,再用Counter统计即可:
# 把所有帖子的名词短语展开成大列表,同样统一小写 all_noun_phrases = [] for phrases in data_noun_phrases: all_noun_phrases.extend([phrase.lower() for phrase in phrases]) # 统计前20个最常见的名词短语 top_noun_phrases = Counter(all_noun_phrases).most_common(20) print("Top 20 常见名词短语:", top_noun_phrases)
小优化:过滤停用词(可选)
如果你的文本里有大量无意义的停用词(比如"the", "a"),可以用NLTK的停用词库过滤掉,让统计结果更精准:
import nltk nltk.download('stopwords') # 第一次使用需要下载停用词库 from nltk.corpus import stopwords stop_words = set(stopwords.words('english')) # 过滤名词里的停用词 nouns = [word.lower() for word, tag in all_tags if tag in noun_tags and word.lower() not in stop_words] # 形容词一般很少有停用词,不过也可以按需过滤 adjectives = [word.lower() for word, tag in all_tags if tag in adj_tags and word.lower() not in stop_words]
这样处理后,得到的高频词汇和短语就更有分析价值啦!
内容的提问来源于stack exchange,提问作者Gioia
相关产品推荐
相关产品推荐

