预处理后DataFrame含列表致split报错,如何修复词数统计函数?
问题描述
初始时,train DataFrame的train列存储文本字符串:
index train 0 My favourite food is anything I didn't have to... 1 Now if he does off himself, everyone will thin... 2 WHY THE FUCK IS BAYLESS ISOING 3 To make her feel threatened 4 Dirty Southern Wankers
原本使用以下函数统计词数:
def word_count(df): word_count = [] for i in df['text']: word = i.split() word_count.append(len(word)) return word_count train['word_count'] = word_count(train)
完成文本预处理后,train列变为单词列表类型:
index train 0 [favourit, food, anyth, didnt, cook] 1 [everyon, think, he, laugh, screw, peopl, inst...] 2 [fuck, bayless, iso] 3 [make, feel, threaten] 4 [dirti, southern, wanker]
此时调用原word_count函数会触发错误:AttributeError: 'list' object has no attribute 'split',原因是当前train列的元素是列表,不再是字符串,无法调用split()方法。
解决方法
方法1:修改原word_count函数
直接取列表的长度即可,无需再执行split,同时修正原函数里列名的笔误:
def word_count(df): word_count = [] for i in df['train']: word_count.append(len(i)) return word_count train['word_count'] = word_count(train)
方法2:用pandas.apply简化实现
不用写循环函数,直接用apply对每个列表取长度:
train['word_count'] = train['train'].apply(len)
方法3:用pandas字符串访问器(更简洁高效)
pandas的str访问器支持对列表类型元素调用len,一行搞定:
train['word_count'] = train['train'].str.len()
内容的提问来源于stack exchange,提问作者Chubiblan
相关产品推荐
相关产品推荐

