You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Atoti中读取DataFrame遇ArrowInvalid错误,如何解决?

解决Atoti session.read_pandas()的ArrowInvalid类型转换错误

这个错误的核心原因是:你指定为键的Row ID列中存在字符串格式的数值(比如示例中的'2531'),但Atoti默认会尝试将键列解析为整数类型,类型不匹配导致转换失败。以下是几种可行的解决方法:

  • 方法一:提前转换DataFrame列类型
    在调用read_pandas()前,先把Row ID列转换为整数类型。注意先处理列中可能存在的非数值字符串,避免转换报错:

    # 转换列类型,无法转换的值设为NaN
    df['Row ID'] = pd.to_numeric(df['Row ID'], errors='coerce')
    # 可选:删除转换后为NaN的无效行
    df = df.dropna(subset=['Row ID'])
    # 转为int64类型
    df['Row ID'] = df['Row ID'].astype('int64')
    # 再执行导入
    global_data = session.read_pandas(df, keys=["Row ID"], table_name="Global_Superstore")
    
  • 方法二:显式指定Atoti的列类型
    通过types参数告诉Atoti将Row ID列按字符串类型处理,跳过自动类型转换:

    from atoti import Type
    global_data = session.read_pandas(
        df,
        keys=["Row ID"],
        table_name="Global_Superstore",
        types={"Row ID": Type.STRING}
    )
    
  • 方法三:清理数据源中的异常值
    先排查Row ID列的内容,过滤掉非纯数字的异常字符串后再导入:

    # 仅保留列值为纯数字的行
    df = df[df['Row ID'].str.match(r'^\d+$')]
    # 转换为int类型
    df['Row ID'] = df['Row ID'].astype('int64')
    # 导入Atoti
    global_data = session.read_pandas(df, keys=["Row ID"], table_name="Global_Superstore")
    

内容的提问来源于stack exchange,提问作者Ardhana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 08:02:07