You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Dask合并含NA键时报错,Pandas merge可正常执行

Dask右连接时因NAType报错的原因分析

问题场景

使用Pandas执行左连接可正常完成,但改用Dask执行对应逻辑的右连接时,抛出TypeError,提示无法将NAType转换为整数。

Pandas正常执行的代码

import pandas as pd

tdf1 = pd.DataFrame([{"id": 1, "val": 4}, {"id": 2, "val": 5}, {"id": 3, "val": 6}, {"id": pd.NA, "val": 7}, {"id": 4, "val": 8}])
tdf2 = pd.DataFrame([{"some_id": 1, "name": "Josh"}, {"some_id": 3, "name": "Jake"}])

# Pandas左连接正常运行
pd.merge(tdf1, tdf2, how="left", left_on="id", right_on="some_id").head()

Dask报错的代码

import dask.dataframe as dd

dd_tdf1 = dd.from_pandas(tdf1, npartitions=10)
dd_tdf2 = dd.from_pandas(tdf2, npartitions=10)

# Dask右连接抛出错误
dd_tdf2.merge(dd_tdf1, left_on="some_id", right_on="id", how="right", npartitions=10).compute(scheduler="threads").head()

报错信息

File /opt/conda/lib/python3.10/site-packages/pandas/core/reshape/merge.py:1585, in <genexpr>(.0)
   1581         return _get_no_sort_one_missing_indexer(left_n, False)
   1583 # get left & right join labels and num. of levels at each location
   1584 mapped = (
-> 1585     _factorize_keys(left_keys[n], right_keys[n], sort=sort, how=how)
   1586     for n in range(len(left_keys))
   1587 )
   1588 zipped = zip(*mapped)
   1589 llab, rlab, shape = (list(x) for x in zipped)

File /opt/conda/lib/python3.10/site-packages/pandas/core/reshape/merge.py:2313, in _factorize_keys(lk, rk, sort, how)
   2309 if is_integer_dtype(lk.dtype) and is_integer_dtype(rk.dtype):
   2310     # GH#23917 TODO: needs tests for case where lk is integer-dtype
   2311     #  and rk is datetime-dtype
   2312     klass = libhashtable.Int64Factorizer
-> 2313     lk = ensure_int64(np.asarray(lk))
   2314     rk = ensure_int64(np.asarray(rk))
   2316 elif needs_i8_conversion(lk.dtype) and is_dtype_equal(lk.dtype, rk.dtype):
   2317     # GH#23917 TODO: Needs tests for non-matching dtypes

File pandas/_libs/algos_common_helper.pxi:86, in pandas._libs.algos.ensure_int64()

TypeError: int() argument must be a string, a bytes-like object or a real number, not 'NAType'

报错原因

  1. 键的类型不匹配与NAType处理漏洞:
    • tdf1['id']是Pandas的可空整数类型Int64,包含pd.NA;而tdf2['some_id']是原生非空整数类型int64。
    • Dask合并底层依赖Pandas的_factorize_keys函数,该函数判断左右键均为整数类型后,会尝试将列强制转换为int64。但pd.NA属于NAType,无法直接转为原生int,因此触发类型错误。
  2. 连接方向触发的逻辑分支差异:
    • Pandas左连接时,对右表键的类型兼容逻辑更灵活;而Dask右连接时,触发了更严格的整数类型转换分支,未考虑可空整数中存在NAType的场景。

解决方法

  • 统一键的可空整数类型:将tdf2['some_id']转为Int64类型,确保左右键类型一致:
    tdf2['some_id'] = tdf2['some_id'].astype('Int64')
    dd_tdf2 = dd.from_pandas(tdf2, npartitions=10)
    # 重新执行合并即可正常运行
    result = dd_tdf2.merge(dd_tdf1, left_on="some_id", right_on="id", how="right", npartitions=10).compute(scheduler="threads")
    print(result.head())
    
  • 过滤空值后合并:如果业务允许,先移除tdf1中id为pd.NA的行:
    tdf1_clean = tdf1.dropna(subset=['id'])
    dd_tdf1_clean = dd.from_pandas(tdf1_clean, npartitions=10)
    result = dd_tdf2.merge(dd_tdf1_clean, left_on="some_id", right_on="id", how="right", npartitions=10).compute(scheduler="threads")
    print(result.head())
    

内容的提问来源于stack exchange,提问作者Jorge Cespedes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 20:10:31