BigQuery调用load_table_from_dataframe触发PyArrow类型错误如何解决
解决PyArrow TypeError:调用BigQuery load_table_from_dataframe时整数预期却得到字符串的问题
问题根因
表层日志里的类型匹配是误判,实际上BigQuery schema中声明的DATE、TIME类型,不能直接对应Pandas的object(字符串)类型。PyArrow在做格式转换时,需要接收对应格式的日期/时间类型值,而非原始字符串,因此抛出类型错误。你之前将时间字段转为秒数的方案不生效,是因为PyArrow的TIME类型对应Python的datetime.time对象,而非整数时间戳。
解决方案
方案1:显式转换DataFrame对应字段为匹配类型
对date_key、start_time_key、end_time_key三个字段做类型转换后再上传,代码如下:
import pandas as pd # 转换日期字段为date类型 df_data["date_key"] = pd.to_datetime(df_data["date_key"]).dt.date # 转换时间字段为time类型 df_data["start_time_key"] = pd.to_datetime(df_data["start_time_key"]).dt.time df_data["end_time_key"] = pd.to_datetime(df_data["end_time_key"]).dt.time
方案2:绕过PyArrow转换,使用CSV作为中间上传格式
如果不想修改原始DataFrame的字段类型,可以在任务配置中指定源格式为CSV,由BigQuery服务端自动完成字符串到对应类型的解析:
job_config = bigquery.job.LoadJobConfig( schema=schema, source_format=bigquery.SourceFormat.CSV )
通用错误字段定位方法
如果后续遇到同类无法定位错误字段的问题,可以逐字段做Arrow转换测试,快速定位不匹配字段:
import pyarrow from google.cloud.bigquery import _pandas_helpers for field in schema: try: series = df_data[field["name"]] arrow_type = _pandas_helpers.bq_field_to_arrow_type(field) pyarrow.Array.from_pandas(series, type=arrow_type) except Exception as e: print(f"字段{field['name']}类型不匹配,错误:{e}")
内容的提问来源于stack exchange,提问作者Philippe Hebert
相关产品推荐
相关产品推荐

