You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python运行pandas分层采样报ValueError: unknown type object解决方案

问题根因

你遇到的ValueError: unknown type object报错触发自numexpr库的类型校验逻辑:pandas的query()方法默认调用numexpr作为执行引擎,numexpr仅支持原生数值、字符串类型,不支持pandas的object dtype列参与运算。你从JSON字段展开得到的分层列虽然手动做了float转int的操作,但仍存在部分列因JSON解析结果类型不一致,最终dtype仍为object的情况,触发了numexpr的类型报错。此前测试正常是因为旧测试数据的分层列类型均为numexpr支持的原生类型,新测试数据展开后存在object类型列导致报错。

修复方案
  • 方案1(快速生效,改动最小):修改自定义stratified_sample函数中所有df.query(qry)调用,显式指定使用Python引擎绕过numexpr校验,改动如下:
    把函数里两处df.query(qry)改为df.query(qry, engine='python')即可。
  • 方案2(彻底修复类型问题):先校验所有分层列的类型,显式转换为int类型,执行如下代码:
    # 检查所有分层列的dtype
    print(df[test].dtypes)
    # 批量转换所有分层列为int类型
    for col in test:
        df[col] = df[col].astype(int)
    
  • 方案3(最稳定,规避字符串拼接和引擎问题):替换动态拼接query的逻辑为原生布尔索引筛选,无需依赖query引擎,也不会出现字符串转义bug,修改stratified_sample中筛选逻辑:
    把原有拼接qry字符串的部分替换为构造布尔掩码:
    # 替换原有qry拼接逻辑
    mask = pd.Series([True]*len(df), index=df.index)
    for s in range(len(strata)):
        stratum = strata[s]
        value = tmp_grpd.iloc[i][stratum]
        mask = mask & (df[stratum] == value)
    # 后续采样改为用mask筛选
    if first:
        stratified_df = df[mask].sample(n=n, random_state=seed).reset_index(drop=(not keep_index))
        first = False
    else:
        tmp_df = df[mask].sample(n=n, random_state=seed).reset_index(drop=(not keep_index))
        stratified_df = stratified_df.append(tmp_df, ignore_index=True)
    

内容的提问来源于stack exchange,提问作者PedroSPSantos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 15:27:03