You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何移除数据集中仅出现一次的元组内词语?

移除数据集中仅出现一次的词语

实现思路

  • 先统计整个数据集中所有词语的全局出现频次
  • 遍历每行数据,仅保留频次大于1的词语

代码实现(Python + Pandas)

import pandas as pd
from collections import Counter

# 加载数据集(示例数据演示,实际替换为你的数据源)
df = pd.DataFrame({
    'before_cleaning': [['cool'], ['gooooood'], ['we', 'love', 'it', 'cool'], ['love', 'it']]
})

# 统计所有词语的全局出现频次
all_words = [word for sublist in df['before_cleaning'] for word in sublist]
word_freq = Counter(all_words)

# 定义过滤函数:移除仅出现一次的词语
def filter_single_occurence_words(word_list):
    return [word for word in word_list if word_freq[word] > 1]

# 应用过滤逻辑到每行数据
df['after_cleaning'] = df['before_cleaning'].apply(filter_single_occurence_words)

# 输出结果
print(df)

输出结果

before_cleaning    after_cleaning
0                [cool]            [cool]
1            [gooooood]                []
2  [we, love, it, cool]  [love, it, cool]
3            [love, it]        [love, it]

说明

  • 该方案对数千行的数据集适配性强,统计和过滤均为线性时间复杂度,处理效率稳定
  • 若你的数据集不是Pandas DataFrame格式,仅需调整数据加载部分,核心的频次统计与过滤逻辑可直接复用

内容的提问来源于stack exchange,提问作者Dewani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 09:20:23