You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何多进程运算比Pandas简单计算更慢?如何并行化模糊字符串比对?

关于多进程比Pandas简单计算慢&并行化模糊字符串比对的解答

一、为什么多进程运算比Pandas简单计算更慢?

这其实是很常见的情况,核心原因在于多进程的额外开销抵消了并行计算的收益,具体来说:

  • 进程本身的开销大:每个进程都需要独立的内存空间,启动、销毁进程以及进程间通信都要消耗时间。而Pandas的简单计算(比如列运算、基础聚合)大多是底层用C实现的向量化操作,单进程下就能跑到接近硬件极限的速度,多进程的额外成本反而拖了后腿。
  • 数据拆分与合并的成本:要并行就得把数据拆分成多个分片分给不同进程,处理完再合并结果。如果数据量不大,拆分合并的时间甚至会超过实际计算的时间,完全得不偿失。
  • Pandas原生操作的优势:很多Pandas操作本身已经绕过了Python的GIL(全局解释器锁),用C直接执行,单进程下就能高效利用CPU资源,多进程并行的提升空间非常有限。

举个例子,你用df['col'] * 2这种向量化操作,底层是C循环直接处理整个列,比拆成多个进程去处理每个小片段快得多。

二、如何用Pandas的apply并行化大量模糊字符串比对?

模糊字符串比对(比如用fuzzywuzzy)属于逐行的非向量化操作,单进程apply速度会很慢,这时候并行化的收益就很明显了。给你三个实用的实现方法:

方法1:用Dask实现并行(适配大数据集)

Dask可以把Pandas的操作自动拆分成并行任务,适合处理大规模数据:

import dask.dataframe as dd
from fuzzywuzzy import fuzz
import pandas as pd

# 构造示例数据
master = pd.DataFrame({
    'original': ['this is a nice sentence', 'this is another one', 'stackoverflow is nice']
})
slave = pd.DataFrame({
    'name': ['hello world', 'congratulations', 'this is a nice sentence ', 'this is another one', 'stackoverflow is awesome']
})

# 将Pandas DataFrame转为Dask DataFrame,分区数建议等于CPU核心数
dask_slave = dd.from_pandas(slave, npartitions=4)

# 定义模糊匹配函数:找到master中匹配度最高的条目和分数
def fuzzy_match(name):
    max_score = 0
    best_match = None
    cleaned_name = name.strip()
    for orig in master['original']:
        score = fuzz.ratio(cleaned_name, orig.strip())
        if score > max_score:
            max_score = score
            best_match = orig
    return (best_match, max_score)

# 并行执行apply,meta参数指定返回结果的类型
result = dask_slave['name'].apply(
    fuzzy_match,
    meta=('result', 'object')  # 因为返回元组,所以用object类型
).compute()

# 将结果合并回原DataFrame
slave[['best_match', 'score']] = pd.DataFrame(result.tolist(), index=slave.index)
print(slave)

方法2:用Swifter自动适配并行(省心首选)

Swifter会自动判断你的apply操作是否适合并行,适合的话就用Dask或多进程加速,不用手动处理细节:

import swifter
import pandas as pd
from fuzzywuzzy import fuzz

master = pd.DataFrame({
    'original': ['this is a nice sentence', 'this is another one', 'stackoverflow is nice']
})
slave = pd.DataFrame({
    'name': ['hello world', 'congratulations', 'this is a nice sentence ', 'this is another one', 'stackoverflow is awesome']
})

# 定义模糊匹配函数,返回Series方便合并
def fuzzy_match(name):
    max_score = 0
    best_match = None
    cleaned_name = name.strip()
    for orig in master['original']:
        score = fuzz.ratio(cleaned_name, orig.strip())
        if score > max_score:
            max_score = score
            best_match = orig
    return pd.Series([best_match, max_score], index=['best_match', 'score'])

# 直接用swifter.apply,自动处理并行逻辑
slave[['best_match', 'score']] = slave['name'].swifter.apply(fuzzy_match)
print(slave)

方法3:用multiprocessing手动实现并行(无依赖方案)

如果不想用第三方库,用Python标准库的multiprocessing也能实现:

import pandas as pd
from fuzzywuzzy import fuzz
from multiprocessing import Pool, cpu_count

master = pd.DataFrame({
    'original': ['this is a nice sentence', 'this is another one', 'stackoverflow is nice']
})
slave = pd.DataFrame({
    'name': ['hello world', 'congratulations', 'this is a nice sentence ', 'this is another one', 'stackoverflow is awesome']
})

# 定义模糊匹配函数
def fuzzy_match(name):
    max_score = 0
    best_match = None
    cleaned_name = name.strip()
    for orig in master['original']:
        score = fuzz.ratio(cleaned_name, orig.strip())
        if score > max_score:
            max_score = score
            best_match = orig
    return (best_match, max_score)

# 主进程中启动进程池并行处理
if __name__ == '__main__':
    # 进程数设为CPU核心数
    with Pool(cpu_count()) as pool:
        results = pool.map(fuzzy_match, slave['name'].tolist())
    
    # 将结果合并回原DataFrame
    slave[['best_match', 'score']] = pd.DataFrame(results, index=slave.index)
    print(slave)

额外优化建议

  • 安装python-Levenshtein库:fuzzywuzzy默认用纯Python实现匹配算法,安装这个库后会换成C实现的Levenshtein算法,速度能提升好几倍,只需要执行pip install python-Levenshtein即可,代码无需修改。
  • 避免在匹配函数中引用超大全局变量:如果master数据集很大,建议把它传入函数或者用其他方式优化,避免进程间的数据拷贝开销。

内容的提问来源于stack exchange,提问作者ℕʘʘḆḽḘ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:54:29