基于元素数量筛选元组:NLP任务中DataFrame数据过滤求助
NLP任务数据筛选:保留词汇数超4的样本
需求说明
正在开展NLP任务,已完成Tokenization(分词)处理,数据已转换为元组形式,需筛选出包含4个以上词汇的样本。
原始数据集示例
ID content 0 [yes, no, check, sample, word] 1 [never, you] 2 [non, program, more, link, draft, ask] 3 [able] 4 [to, ask, you, other, man, will]
期望结果
仅保留词汇数超过4的样本(ID为0、2、4):
ID content 0 [yes, no, check, sample, word] 2 [non, program, more, link, draft, ask] 4 [to, ask, you, other, man, will]
代码修正
你当前的代码存在两处问题:
ne()是Pandas中的“不等于”判断方法,不符合“筛选长度大于4”的需求;ne(>4)语法不合法,无法将比较运算符直接作为参数传入方法。
正确写法1(直接用比较运算符)
df_new = df.loc[df.content.map(len) > 4]
正确写法2(使用Pandas的gt()方法,即greater than)
df_new = df.loc[df.content.map(len).gt(4)]
以上两种写法都能正确筛选出content列中元素数量超过4的样本,得到你期望的数据集。
内容的提问来源于stack exchange,提问作者Dewani
相关产品推荐
相关产品推荐

