You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于元素数量筛选元组:NLP任务中DataFrame数据过滤求助

NLP任务数据筛选:保留词汇数超4的样本

需求说明

正在开展NLP任务,已完成Tokenization(分词)处理,数据已转换为元组形式,需筛选出包含4个以上词汇的样本。

原始数据集示例

ID                                content
 0         [yes, no, check, sample, word]
 1                           [never, you]
 2 [non, program, more, link, draft, ask]
 3                                 [able]
 4       [to, ask, you, other, man, will]

期望结果

仅保留词汇数超过4的样本(ID为0、2、4):

ID                                content
 0         [yes, no, check, sample, word]
 2 [non, program, more, link, draft, ask]
 4       [to, ask, you, other, man, will]

代码修正

你当前的代码存在两处问题:

  1. ne()是Pandas中的“不等于”判断方法,不符合“筛选长度大于4”的需求;
  2. ne(>4)语法不合法,无法将比较运算符直接作为参数传入方法。

正确写法1(直接用比较运算符)

df_new = df.loc[df.content.map(len) > 4]

正确写法2(使用Pandas的gt()方法,即greater than)

df_new = df.loc[df.content.map(len).gt(4)]

以上两种写法都能正确筛选出content列中元素数量超过4的样本,得到你期望的数据集。

内容的提问来源于stack exchange,提问作者Dewani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 16:40:16