You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升Pandas DataFrame中列表推导式的运行速度?大数据集下除列表推导式外是否有更快的集合元素过滤方法

嘿,针对你提出的两个Pandas性能优化问题,我来分享些实用的解决方案,都是处理大数据集时亲测有效的技巧:

1. 提升Pandas DataFrame中列表推导式的运行速度

你已经把停用词转成集合的操作非常正确——集合的in操作是O(1)时间复杂度,比列表快太多,这点一定要保留。在此基础上,还可以通过以下方式进一步提速:

  • 用Pandas内置的向量化方法替代纯Python循环
    你当前用的apply(lambda x: ...)本质是在做Python级别的循环,而Pandas的字符串方法是底层C优化的,速度会快很多。可以把拆分、过滤、合并改成链式的向量化操作:

    list_stopwords = set(stop_words.get_stop_words('en'))
    # 拆分字符串为单词列表 → 过滤停用词 → 重新合并为字符串
    data['description'] = data['description'].str.split() \
      .apply(lambda x: [word for word in x if word not in list_stopwords]) \
      .str.join(' ')
    

    这里的str.split()和str.join()都是向量化实现,比纯Python的split()和join()在批量处理时效率更高。

  • 借助swifter库自动选择最优执行方式
    swifter会自动判断你的操作适合用向量化还是并行处理,无需手动调整,对新手非常友好:
    先安装库:

    pip install swifter
    

    然后修改代码:

    import swifter
    list_stopwords = set(stop_words.get_stop_words('en'))
    data['description'] = data['description'].swifter.apply(
        lambda x: " ".join([word for word in x.split() if word not in list_stopwords])
    )
    
  • 统一字符串大小写(可选)
    如果你的停用词都是小写,建议先把description转成小写,避免因大小写不匹配导致的无效过滤,同时也能减少判断逻辑:

    data['description'] = data['description'].str.lower().swifter.apply(
        lambda x: " ".join([word for word in x.split() if word not in list_stopwords])
    )
    
2. 大数据集下更快的集合元素过滤方法

当数据集规模很大时,apply+列表推导的效率还是会受限,推荐以下几种更高效的方案:

  • 向量化拆分+过滤+聚合
    这种方法完全利用Pandas的分组和聚合能力,避免Python循环:

    list_stopwords = set(stop_words.get_stop_words('en'))
    # 1. 把每个字符串拆成单词列表,然后展开成每行一个单词
    exploded_data = data.assign(words=data['description'].str.split()).explode('words')
    # 2. 过滤掉属于停用词的行
    filtered_data = exploded_data[~exploded_data['words'].isin(list_stopwords)]
    # 3. 按原数据的索引分组,重新合并单词为字符串
    data['description'] = filtered_data.groupby(filtered_data.index)['words'].str.join(' ')
    

    这种方式的核心是用Pandas的内置函数替代Python循环,底层是C实现,速度能提升数倍甚至数十倍。

  • 用NumPy加速过滤逻辑
    NumPy的数组操作比纯Python列表更快,可以把单词列表转成NumPy数组后再过滤:

    import numpy as np
    stopwords_np = np.array(list(stop_words.get_stop_words('en')))
    
    def filter_with_numpy(text):
        words = np.array(text.split())
        # 生成过滤掩码:保留不在停用词里的单词
        keep_mask = ~np.in1d(words, stopwords_np)
        return ' '.join(words[keep_mask])
    
    data['description'] = data['description'].apply(filter_with_numpy)
    
  • 并行处理超大规模数据集
    如果你的数据集大到单进程处理太慢,可以用Dask实现分块并行处理:

    import dask.dataframe as dd
    list_stopwords = set(stop_words.get_stop_words('en'))
    # 把Pandas DataFrame转成Dask DataFrame,根据CPU核心数设置分区数
    dask_df = dd.from_pandas(data, npartitions=4)
    # 定义处理函数,注意指定meta参数告诉Dask返回的数据类型
    def clean_description(text):
        return " ".join([word for word in text.split() if word not in list_stopwords])
    
    dask_df['description'] = dask_df['description'].apply(clean_description, meta=('description', 'object'))
    # 计算结果并转回Pandas DataFrame
    data = dask_df.compute()
    
  • 正则表达式批量替换(适合停用词数量较少的场景)
    构建匹配整个单词的正则表达式,一次性替换所有停用词,速度非常快,但要注意处理标点和边界问题:

    import re
    list_stopwords = set(stop_words.get_stop_words('en'))
    # 构建匹配整个单词的正则模式,转义特殊字符避免匹配错误
    stopword_regex = r'\b(' + '|'.join(re.escape(word) for word in list_stopwords) + r')\b'
    # 替换停用词并清理多余空格
    data['description'] = data['description'].str.replace(stopword_regex, '', regex=True)
    data['description'] = data['description'].str.replace(r'\s+', ' ', regex=True).str.strip()
    

总的来说,优先尝试向量化操作和swifter,如果还是不够快,再考虑并行处理或者正则替换,根据你的数据集规模和停用词数量选择最合适的方案。

内容的提问来源于stack exchange,提问作者chuky pedro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 14:12:34