You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何加速Pandas Categorical类创建?解决十亿级行数据性能瓶颈

十亿行DataFrame转Categorical的性能优化方案

针对你遇到的pd.Categorical初始化时coerce_indexer_dtype耗时问题,结合你已经提前格式化数据的前提,给你两个高效解决方案:

方案一:使用官方API pd.Categorical.from_codes(优先推荐)

既然你的数据已经是与categories完全匹配的索引值(即合法的codes),直接用from_codes构造Categorical是最安全且高效的方式——这个方法专门为直接传入codes的场景设计,内部不会执行coerce_indexer_dtype这类冗余检查。

示例代码:

import pandas as pd
import numpy as np

# 替换为你的十亿级numpy数组
arr0 = np.tile(np.arange(1000, dtype=np.int16), 1000)
arr1 = np.repeat(np.arange(1000, dtype=np.int16), 1000)

cats = np.arange(1000, dtype=np.int16)
# 直接通过from_codes创建Categorical
cat0 = pd.Categorical.from_codes(arr0, categories=cats, ordered=True)
cat1 = pd.Categorical.from_codes(arr1, categories=cats, ordered=True)

# 构造DataFrame
df = pd.DataFrame({0: cat0, 1: cat1})

方案二:直接操作Categorical内部属性(性能极限优化)

如果追求极致性能,可以绕过pd.Categorical的__init__逻辑,直接赋值内部的_codes和_data属性。这种方式完全跳过所有初始化检查,但依赖pandas内部实现,版本更新后可能需要调整。

示例代码:

import pandas as pd
import numpy as np

arr0 = np.tile(np.arange(1000, dtype=np.int16), 1000)
arr1 = np.repeat(np.arange(1000, dtype=np.int16), 1000)

cats = np.arange(1000, dtype=np.int16)
dtype = pd.CategoricalDtype(categories=cats, ordered=True)

# 初始化空Categorical后直接赋值内部属性
cat0 = pd.Categorical([], dtype=dtype)
cat0._codes = arr0
cat0._data = pd.core.arrays.categorical.CategoricalData(
    categories=cats, ordered=True, fastpath=True
)

cat1 = pd.Categorical([], dtype=dtype)
cat1._codes = arr1
cat1._data = pd.core.arrays.categorical.CategoricalData(
    categories=cats, ordered=True, fastpath=True
)

df = pd.DataFrame({0: cat0, 1: cat1})

补充说明

你之前用fastpath=True仍有性能瓶颈,是因为即便开启fastpath,pd.Categorical的__init__还是会执行coerce_indexer_dtype来确保codes的类型合规;而上述两种方法都彻底跳过了这个步骤,因此能大幅提升十亿级数据的处理速度。

内容的提问来源于stack exchange,提问作者boxblox

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 13:01:03