为何多进程运算比Pandas简单计算更慢?如何并行化模糊字符串比对?
关于多进程比Pandas简单计算慢&并行化模糊字符串比对的解答
一、为什么多进程运算比Pandas简单计算更慢?
这其实是很常见的情况,核心原因在于多进程的额外开销抵消了并行计算的收益,具体来说:
- 进程本身的开销大:每个进程都需要独立的内存空间,启动、销毁进程以及进程间通信都要消耗时间。而Pandas的简单计算(比如列运算、基础聚合)大多是底层用C实现的向量化操作,单进程下就能跑到接近硬件极限的速度,多进程的额外成本反而拖了后腿。
- 数据拆分与合并的成本:要并行就得把数据拆分成多个分片分给不同进程,处理完再合并结果。如果数据量不大,拆分合并的时间甚至会超过实际计算的时间,完全得不偿失。
- Pandas原生操作的优势:很多Pandas操作本身已经绕过了Python的GIL(全局解释器锁),用C直接执行,单进程下就能高效利用CPU资源,多进程并行的提升空间非常有限。
举个例子,你用df['col'] * 2这种向量化操作,底层是C循环直接处理整个列,比拆成多个进程去处理每个小片段快得多。
二、如何用Pandas的apply并行化大量模糊字符串比对?
模糊字符串比对(比如用fuzzywuzzy)属于逐行的非向量化操作,单进程apply速度会很慢,这时候并行化的收益就很明显了。给你三个实用的实现方法:
方法1:用Dask实现并行(适配大数据集)
Dask可以把Pandas的操作自动拆分成并行任务,适合处理大规模数据:
import dask.dataframe as dd from fuzzywuzzy import fuzz import pandas as pd # 构造示例数据 master = pd.DataFrame({ 'original': ['this is a nice sentence', 'this is another one', 'stackoverflow is nice'] }) slave = pd.DataFrame({ 'name': ['hello world', 'congratulations', 'this is a nice sentence ', 'this is another one', 'stackoverflow is awesome'] }) # 将Pandas DataFrame转为Dask DataFrame,分区数建议等于CPU核心数 dask_slave = dd.from_pandas(slave, npartitions=4) # 定义模糊匹配函数:找到master中匹配度最高的条目和分数 def fuzzy_match(name): max_score = 0 best_match = None cleaned_name = name.strip() for orig in master['original']: score = fuzz.ratio(cleaned_name, orig.strip()) if score > max_score: max_score = score best_match = orig return (best_match, max_score) # 并行执行apply,meta参数指定返回结果的类型 result = dask_slave['name'].apply( fuzzy_match, meta=('result', 'object') # 因为返回元组,所以用object类型 ).compute() # 将结果合并回原DataFrame slave[['best_match', 'score']] = pd.DataFrame(result.tolist(), index=slave.index) print(slave)
方法2:用Swifter自动适配并行(省心首选)
Swifter会自动判断你的apply操作是否适合并行,适合的话就用Dask或多进程加速,不用手动处理细节:
import swifter import pandas as pd from fuzzywuzzy import fuzz master = pd.DataFrame({ 'original': ['this is a nice sentence', 'this is another one', 'stackoverflow is nice'] }) slave = pd.DataFrame({ 'name': ['hello world', 'congratulations', 'this is a nice sentence ', 'this is another one', 'stackoverflow is awesome'] }) # 定义模糊匹配函数,返回Series方便合并 def fuzzy_match(name): max_score = 0 best_match = None cleaned_name = name.strip() for orig in master['original']: score = fuzz.ratio(cleaned_name, orig.strip()) if score > max_score: max_score = score best_match = orig return pd.Series([best_match, max_score], index=['best_match', 'score']) # 直接用swifter.apply,自动处理并行逻辑 slave[['best_match', 'score']] = slave['name'].swifter.apply(fuzzy_match) print(slave)
方法3:用multiprocessing手动实现并行(无依赖方案)
如果不想用第三方库,用Python标准库的multiprocessing也能实现:
import pandas as pd from fuzzywuzzy import fuzz from multiprocessing import Pool, cpu_count master = pd.DataFrame({ 'original': ['this is a nice sentence', 'this is another one', 'stackoverflow is nice'] }) slave = pd.DataFrame({ 'name': ['hello world', 'congratulations', 'this is a nice sentence ', 'this is another one', 'stackoverflow is awesome'] }) # 定义模糊匹配函数 def fuzzy_match(name): max_score = 0 best_match = None cleaned_name = name.strip() for orig in master['original']: score = fuzz.ratio(cleaned_name, orig.strip()) if score > max_score: max_score = score best_match = orig return (best_match, max_score) # 主进程中启动进程池并行处理 if __name__ == '__main__': # 进程数设为CPU核心数 with Pool(cpu_count()) as pool: results = pool.map(fuzzy_match, slave['name'].tolist()) # 将结果合并回原DataFrame slave[['best_match', 'score']] = pd.DataFrame(results, index=slave.index) print(slave)
额外优化建议
- 安装
python-Levenshtein库:fuzzywuzzy默认用纯Python实现匹配算法,安装这个库后会换成C实现的Levenshtein算法,速度能提升好几倍,只需要执行pip install python-Levenshtein即可,代码无需修改。 - 避免在匹配函数中引用超大全局变量:如果master数据集很大,建议把它传入函数或者用其他方式优化,避免进程间的数据拷贝开销。
内容的提问来源于stack exchange,提问作者ℕʘʘḆḽḘ
相关产品推荐
相关产品推荐

