You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决Python Pandas使用assign创建新列时出现的TypeError报错

错误原因

你的写法核心问题是sent_tokenize只接收单个字符串作为入参,但你直接传入的x['description']/df['description']是pandas的Series对象,不是单个字符串,所以才抛出类型错误,和列的dtype是object没有关系,pandas的字符串列默认就是object类型,你确认每个元素都是字符串就不存在类型问题。

正确写法

用Series.apply()将sent_tokenize逐行作用到每个description字符串上即可,完全可以在assign中实现,不影响链式操作:

import nltk
# 首次运行需要下载punkt分词模型
# nltk.download('punkt')
from nltk.tokenize import sent_tokenize

df = df.assign(
    sentence_count = lambda x: x['description'].apply(lambda s: len(sent_tokenize(s)))
)

如果要提升大数据量下的运行效率,也可以用列表推导式替代apply,性能会更好:

df = df.assign(
    sentence_count = [len(sent_tokenize(s)) for s in df['description']]
)

内容的提问来源于stack exchange,提问作者mmz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 21:24:01