You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BigQuery中INT64列表UNNEST时JSON序列化错误解决方案问询

解决BigQuery中INT64类型ID列表筛选的JSON序列化错误

最直接的解决方法是将查询参数的类型指定为INT64,而非STRING,无需转换ID列表类型,也不用在查询中对表字段做类型转换,完全匹配BigQuery中t1.id的INT64类型。

修改后的代码如下:

client = bigquery.Client(credentials=credentials)

query = """
SELECT t1.*,
t2.*,
t3.*,
t4.* FROM `<project>.<dataset>.<tabel1>` as t1
join `<project>.<dataset>.<tabel2>` as t2
on t1.label = t2.id
join  `<project>.<dataset>.<tabel3>` as t3
on t3.A = t2.A
join `<project>.<dataset>.<tabel4>` as t4
on t4.obj= t2.obj and t4.A = t3.A
where  t1.id in unnest(@list)
"""
# 将参数类型从"STRING"改为"INT64"
job_config = bigquery.QueryJobConfig(query_parameters=[
                    bigquery.ArrayQueryParameter("list", "INT64", list),
])
choices= client.query(query, job_config=job_config).to_dataframe()

你的ID列表保持原始整数类型即可:

list = [3651056, 3651049, 3640195, 3629411, 3627024,3624939]

错误原因与方案优势

之前报错TypeError: Object of type int64 is not JSON serializable,是因为你把参数类型定义为STRING,但传入了整数列表。BigQuery序列化参数时,整数类型(尤其是numpy int64类型)无法直接转为STRING类型的JSON值。

改用INT64类型参数后:

  1. 彻底避免序列化错误,逻辑更直观
  2. 能利用t1.id字段的索引(如果存在),不会像cast(t1.id as STRING)方案那样触发全表扫描——40亿行的场景下,性能差异会非常显著

内容的提问来源于stack exchange,提问作者Serge de Gosson de Varennes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 16:37:33