You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

向BigQuery推送含列表字段的表时遇ArrowTypeError问题求助

问题解决:本地执行to_gbq报ArrowTypeError,Colab正常运行

问题根源

本地环境(Python3.10 + Pandas1.5.1)中,高版本的pandas-gbq依赖pyarrow进行数据序列化,默认不会自动将Pandas的列表类型Series映射为BigQuery的重复数组类型;而Colab的旧环境(Python3.7 + Pandas1.3.5)使用了旧版序列化逻辑,隐式处理了列表类型的转换,因此可以正常运行。

解决方案

方案1:显式指定BigQuery表架构(推荐)

直接定义表的字段类型,告诉pyarrow将列A识别为重复整数类型:

import pandas as pd
from google.oauth2 import service_account
from google.cloud.bigquery import SchemaField

df = pd.DataFrame()
df['A'] = [[1], [2], [3]]

credentials = service_account.Credentials.from_service_account_info({--credential infos--})

# 定义表架构,指定A列为重复整数类型
schema = [
    SchemaField('A', 'INTEGER', mode='REPEATED')
]

df.to_gbq(
    destination_table='raw.test',
    project_id='project-test',
    credentials=credentials,
    if_exists='replace',
    table_schema=schema
)

方案2:临时转换列表为字符串(应急用)

如果不需要严格的数组类型,可以先将列表转为字符串,写入BigQuery后再在SQL中转换回来:

import pandas as pd
from google.oauth2 import service_account

df = pd.DataFrame()
df['A'] = [[1], [2], [3]]
# 将列表转为字符串
df['A'] = df['A'].apply(str)

credentials = service_account.Credentials.from_service_account_info({--credential infos--})

df.to_gbq(
    destination_table='raw.test',
    project_id='project-test',
    credentials=credentials,
    if_exists='replace'
)

之后在BigQuery中恢复数组类型:

SELECT SAFE.PARSE_JSON(A) AS A FROM raw.test

方案3:降级依赖版本(不推荐)

匹配Colab的环境版本,安装旧版pandas和pandas-gbq,但这种方法会牺牲新特性,且可能带来其他兼容性问题,仅作为最后备选:

pip install pandas==1.3.5 pandas-gbq==0.17.9

内容的提问来源于stack exchange,提问作者Pierre

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 19:50:36