如何移除数据集中仅出现一次的元组内词语?
移除数据集中仅出现一次的词语
实现思路
- 先统计整个数据集中所有词语的全局出现频次
- 遍历每行数据,仅保留频次大于1的词语
代码实现(Python + Pandas)
import pandas as pd from collections import Counter # 加载数据集(示例数据演示,实际替换为你的数据源) df = pd.DataFrame({ 'before_cleaning': [['cool'], ['gooooood'], ['we', 'love', 'it', 'cool'], ['love', 'it']] }) # 统计所有词语的全局出现频次 all_words = [word for sublist in df['before_cleaning'] for word in sublist] word_freq = Counter(all_words) # 定义过滤函数:移除仅出现一次的词语 def filter_single_occurence_words(word_list): return [word for word in word_list if word_freq[word] > 1] # 应用过滤逻辑到每行数据 df['after_cleaning'] = df['before_cleaning'].apply(filter_single_occurence_words) # 输出结果 print(df)
输出结果
before_cleaning after_cleaning 0 [cool] [cool] 1 [gooooood] [] 2 [we, love, it, cool] [love, it, cool] 3 [love, it] [love, it]
说明
- 该方案对数千行的数据集适配性强,统计和过滤均为线性时间复杂度,处理效率稳定
- 若你的数据集不是Pandas DataFrame格式,仅需调整数据加载部分,核心的频次统计与过滤逻辑可直接复用
内容的提问来源于stack exchange,提问作者Dewani
相关产品推荐
相关产品推荐

