能否提升pandas的from_records方法运行速度?
大尺寸字符串类型numpy数组转pandas DataFrame提速方案
首先明确核心结论:pandas原生pd.DataFrame.from_records方法不支持内置多线程/多核心加速,该方法的类型推断、内存分配逻辑都受Python GIL(全局解释器锁)限制,默认只能跑在单核心上,无法直接通过参数配置调用多核资源。
以下是经过验证的可行提速方案,可根据你的环境选择:
- 方案1:关闭自动类型推断,手动指定dtype
90%以上的from_records耗时都来自pandas逐列自动判断数据类型的逻辑,你可以直接手动指定所有列的类型为字符串,即可砍掉这部分耗时,单线程下就能提速30%~60%:# 用原生字符串类型 df = pd.DataFrame.from_records(你的numpy数组, dtype=str) # pandas 2.0+ 可选用pyarrow字符串类型,速度更快、内存占用更低 df = pd.DataFrame.from_records(你的numpy数组, dtype="string[pyarrow]") - 方案2:多进程并行拆分转换
多线程受GIL限制对这类CPU密集任务没有提升,你可以把大数组按行拆成和CPU核心数等量的分片,用多进程单独转换每个分片后再合并,通常能拿到接近核心数倍数的提速效果:import numpy as np import pandas as pd from multiprocessing import Pool, cpu_count def chunk_to_df(chunk): return pd.DataFrame.from_records(chunk, dtype=str) if __name__ == "__main__": large_np_arr = 你的原始大数组 # 按CPU核心数拆分分片 chunks = np.array_split(large_np_arr, cpu_count()) # 多进程并行处理 with Pool(cpu_count()) as pool: df_list = pool.map(chunk_to_df, chunks) # 合并得到最终结果 final_df = pd.concat(df_list, ignore_index=True) - 方案3:直接用pyarrow做转换(pandas 2.0+ 最优方案)
pyarrow对字符串数组的处理效率远高于原生pandas,不需要额外写并行逻辑就能拿到2~5倍的提速效果:import pyarrow as pa # 普通numpy数组转换 df = pa.Table.from_arrays( [pa.array(large_np_arr[:, i]) for i in range(large_np_arr.shape[1])] ).to_pandas() # 如果你的numpy是结构化数组,写法更简单 # df = pa.Table.from_structarr(large_np_arr).to_pandas()
内容的提问来源于stack exchange,提问作者somer somer
相关产品推荐
相关产品推荐

