如何用Pandas高效实现Series每行词频统计(生成字典)
我有一个Pandas DataFrame,其中包含名为spam['v2']的Series,每行都是一个句子。想要生成一个新Series,每行对应一个以单词为键、出现次数为值的字典,用来统计该行的词频。
示例说明
原Series内容
假设原Series的部分行内容如下:
Go until jurong point, crazy.. Available only in bugis n great world la e buffet... Cine there got amore wat...
Ok lar... Joking wif u oni...
Free entry in 2 a wkly comp to win FA Cup final tkts 21st May 2005. Text FA to 87121 to receive entry question(std txt rate)T&C's apply 08452810075over18's
期望生成的新Series
每行对应词频字典,示例如下:
第一行:
{'Go': 1, 'until': 1, 'jurong': 1, 'point,': 1, 'crazy..': 1, ..., 'wat...': 1}
第二行:{'Ok': 1, 'lar...': 1, 'Joking': 1, 'wif': 1, 'u': 1, 'oni...': 1}
第三行:{'Free': 1, 'entry': 2, 'in': 1, '2': 1, ..., 'over18\'s': 1}(其中"entry"出现2次,对应值为2)
现有实现代码
我已经通过自定义Python函数结合apply方法实现了需求,代码如下:
import pandas as pd spam = pd.read_csv('spam.csv') def freq(text): words = text.split() wfreq = [words.count(w) for w in words] return dict(zip(words, wfreq)) count = spam['v2'].apply(freq) count = pd.Series(count)
问题
我希望找到更高效的实现方式,直接利用Pandas或Series的内置方法完成该任务,不需要依赖自定义的Python函数,求具体的实现方案。
内容的提问来源于stack exchange,提问作者Mubashir Ahmed Siddiqui

