You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

预处理后DataFrame含列表致split报错,如何修复词数统计函数?

问题描述

初始时,train DataFrame的train列存储文本字符串:

index                     train
0        My favourite food is anything I didn't have to...
1        Now if he does off himself, everyone will thin...
2                           WHY THE FUCK IS BAYLESS ISOING
3                              To make her feel threatened
4                                   Dirty Southern Wankers

原本使用以下函数统计词数:

def word_count(df):
    word_count = []
    for i in df['text']:
        word = i.split()
        word_count.append(len(word))
    return word_count

train['word_count'] = word_count(train)

完成文本预处理后,train列变为单词列表类型:

index                             train
0                     [favourit, food, anyth, didnt, cook]
1        [everyon, think, he, laugh, screw, peopl, inst...]
2                                     [fuck, bayless, iso]
3                                   [make, feel, threaten]
4                                [dirti, southern, wanker]

此时调用原word_count函数会触发错误:AttributeError: 'list' object has no attribute 'split',原因是当前train列的元素是列表,不再是字符串,无法调用split()方法。

解决方法

方法1:修改原word_count函数

直接取列表的长度即可,无需再执行split,同时修正原函数里列名的笔误:

def word_count(df):
    word_count = []
    for i in df['train']:
        word_count.append(len(i))
    return word_count

train['word_count'] = word_count(train)

方法2:用pandas.apply简化实现

不用写循环函数,直接用apply对每个列表取长度:

train['word_count'] = train['train'].apply(len)

方法3:用pandas字符串访问器(更简洁高效)

pandas的str访问器支持对列表类型元素调用len,一行搞定:

train['word_count'] = train['train'].str.len()

内容的提问来源于stack exchange,提问作者Chubiblan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 21:45:03