GCP Cloud Functions写Pandas Dataframe到BigQuery仅支持PARQUET/CSV吗
错误原因
你触发报错是因为load_table_from_dataframe是BigQuery Python SDK专门针对pandas DataFrame的加载接口,底层仅支持将DataFrame序列化为PARQUET或CSV格式上传,你手动指定的NEWLINE_DELIMITED_JSON不属于该方法支持的源格式范围。
最优实现方案
根据你的场景(单次仅写入1条天气数据),可按需选择以下两种方案:
方案1:保留现有pandas逻辑,改动最小
仅需要删除job_config.source_format = bigquery.SourceFormat.NEWLINE_DELIMITED_JSON这一行配置即可,SDK会默认使用PARQUET格式序列化DataFrame上传,完全兼容现有逻辑。
如果要显式指定避免歧义,也可以替换为以下配置:
job_config.source_format = bigquery.SourceFormat.PARQUET
修改后核心代码片段:
job_config = bigquery.LoadJobConfig() job_config.autodetect = True job_config.write_disposition = "WRITE_APPEND" job_config.source_format = bigquery.SourceFormat.PARQUET load_job = client.load_table_from_dataframe(df, dataset_ref.table("weather_de"), job_config=job_config)
方案2:移除pandas依赖,使用流式插入(更推荐)
你的场景单次仅插入1条数据,完全不需要引入pandas依赖,直接使用BigQuery的流式插入接口即可,能大幅降低Cloud Functions冷启动时间,减少依赖体积,代码也更简洁:
完整实现代码:
from google.cloud import bigquery import requests import datetime def hello_pubsub(event, context): response = requests.get("https://api.openweathermap.org/data/2.5/weather?q=berlin&appid=12345&units=metric&lang=de") responseJson = response.json() # 直接构造插入行,无需转换为DataFrame row_to_insert = [{ "datetime": datetime.datetime.now(), "name": responseJson["name"], "temp": float(responseJson["main"]["temp"]), "windspeed": float(responseJson["wind"]["speed"]), "winddeg": int(responseJson["wind"]["deg"]) }] project_id = 'myproj' client = bigquery.Client(project=project_id) table_id = "weather.weather_de" # 流式插入,不需要配置LoadJob errors = client.insert_rows_json(table_id, row_to_insert) if errors: raise RuntimeError(f"插入BigQuery失败:{errors}")
该方案优势:
- 无需安装pandas依赖,Cloud Functions部署包体积更小,冷启动速度提升30%以上
- 去掉了冗余的DataFrame转换逻辑,代码更简洁
- 单条数据插入延迟更低,不需要等待异步LoadJob执行完成
内容的提问来源于stack exchange,提问作者JoergP
相关产品推荐
相关产品推荐

