Cython与Pandas的ChainedAssignmentError问题:引用计数差异处理
问题详情
使用Cython包加速Pandas操作时,因Cython与标准Python的引用计数差异,触发ChainedAssignmentError的FutureWarning。
Cython化脚本代码:
import cython import pandas as pd import sys def main(): df = pd.DataFrame({"A": [1, 2, 3]}) df["A"] = df["A"].astype(object) return df
触发的警告内容:
FutureWarning: ChainedAssignmentError: behaviour will change in pandas 3.0!
You are setting values through chained assignment. Currently this works in certain cases, but when using Copy-on-Write (which will become the default behaviour in pandas 3.0) this will never work to update the original DataFrame or Series, because the intermediate object on which we are setting values will behave as a copy.
相同代码在标准Python环境运行完全正常,但隐藏警告并非最优方案——Pandas 3.0发布后该代码会直接失效。
问题根源在于Pandas的ChainedAssignment检查逻辑:Pandas通过系统引用数判断对象是否为临时对象,但Cython化代码默认比标准Python少一个引用。可调用sys.getrefcount(df)对比两者的引用数差异验证这一点。目前临时规避方法是在astype前添加references = [df]增加引用数,即可消除警告。
需求
- a. 该问题的内部修复方法(无需依赖临时规避手段)
- b. 向Pandas团队反馈此问题的官方途径(Pandas用户指南明确说明支持Cython)
复现步骤
- 安装Cython
pip install Cython
- 创建
setup.py文件:
from setuptools import setup from Cython.Build import cythonize setup( ext_modules = cythonize("run.pyx") )
- 生成Cython化扩展文件:
python setup.py build_ext --inplace
- 运行测试代码:
import run run.main()
内容的提问来源于stack exchange,提问作者Olusola Nylander

