如何提升Pandas DataFrame中列表推导式的运行速度?大数据集下除列表推导式外是否有更快的集合元素过滤方法
嘿,针对你提出的两个Pandas性能优化问题,我来分享些实用的解决方案,都是处理大数据集时亲测有效的技巧:
你已经把停用词转成集合的操作非常正确——集合的in操作是O(1)时间复杂度,比列表快太多,这点一定要保留。在此基础上,还可以通过以下方式进一步提速:
用Pandas内置的向量化方法替代纯Python循环
你当前用的apply(lambda x: ...)本质是在做Python级别的循环,而Pandas的字符串方法是底层C优化的,速度会快很多。可以把拆分、过滤、合并改成链式的向量化操作:list_stopwords = set(stop_words.get_stop_words('en')) # 拆分字符串为单词列表 → 过滤停用词 → 重新合并为字符串 data['description'] = data['description'].str.split() \ .apply(lambda x: [word for word in x if word not in list_stopwords]) \ .str.join(' ')这里的
str.split()和str.join()都是向量化实现,比纯Python的split()和join()在批量处理时效率更高。借助
swifter库自动选择最优执行方式swifter会自动判断你的操作适合用向量化还是并行处理,无需手动调整,对新手非常友好:
先安装库:pip install swifter然后修改代码:
import swifter list_stopwords = set(stop_words.get_stop_words('en')) data['description'] = data['description'].swifter.apply( lambda x: " ".join([word for word in x.split() if word not in list_stopwords]) )统一字符串大小写(可选)
如果你的停用词都是小写,建议先把description转成小写,避免因大小写不匹配导致的无效过滤,同时也能减少判断逻辑:data['description'] = data['description'].str.lower().swifter.apply( lambda x: " ".join([word for word in x.split() if word not in list_stopwords]) )
当数据集规模很大时,apply+列表推导的效率还是会受限,推荐以下几种更高效的方案:
向量化拆分+过滤+聚合
这种方法完全利用Pandas的分组和聚合能力,避免Python循环:list_stopwords = set(stop_words.get_stop_words('en')) # 1. 把每个字符串拆成单词列表,然后展开成每行一个单词 exploded_data = data.assign(words=data['description'].str.split()).explode('words') # 2. 过滤掉属于停用词的行 filtered_data = exploded_data[~exploded_data['words'].isin(list_stopwords)] # 3. 按原数据的索引分组,重新合并单词为字符串 data['description'] = filtered_data.groupby(filtered_data.index)['words'].str.join(' ')这种方式的核心是用Pandas的内置函数替代Python循环,底层是C实现,速度能提升数倍甚至数十倍。
用NumPy加速过滤逻辑
NumPy的数组操作比纯Python列表更快,可以把单词列表转成NumPy数组后再过滤:import numpy as np stopwords_np = np.array(list(stop_words.get_stop_words('en'))) def filter_with_numpy(text): words = np.array(text.split()) # 生成过滤掩码:保留不在停用词里的单词 keep_mask = ~np.in1d(words, stopwords_np) return ' '.join(words[keep_mask]) data['description'] = data['description'].apply(filter_with_numpy)并行处理超大规模数据集
如果你的数据集大到单进程处理太慢,可以用Dask实现分块并行处理:import dask.dataframe as dd list_stopwords = set(stop_words.get_stop_words('en')) # 把Pandas DataFrame转成Dask DataFrame,根据CPU核心数设置分区数 dask_df = dd.from_pandas(data, npartitions=4) # 定义处理函数,注意指定meta参数告诉Dask返回的数据类型 def clean_description(text): return " ".join([word for word in text.split() if word not in list_stopwords]) dask_df['description'] = dask_df['description'].apply(clean_description, meta=('description', 'object')) # 计算结果并转回Pandas DataFrame data = dask_df.compute()正则表达式批量替换(适合停用词数量较少的场景)
构建匹配整个单词的正则表达式,一次性替换所有停用词,速度非常快,但要注意处理标点和边界问题:import re list_stopwords = set(stop_words.get_stop_words('en')) # 构建匹配整个单词的正则模式,转义特殊字符避免匹配错误 stopword_regex = r'\b(' + '|'.join(re.escape(word) for word in list_stopwords) + r')\b' # 替换停用词并清理多余空格 data['description'] = data['description'].str.replace(stopword_regex, '', regex=True) data['description'] = data['description'].str.replace(r'\s+', ' ', regex=True).str.strip()
总的来说,优先尝试向量化操作和swifter,如果还是不够快,再考虑并行处理或者正则替换,根据你的数据集规模和停用词数量选择最合适的方案。
内容的提问来源于stack exchange,提问作者chuky pedro

