You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Polars DataFrame写入BigQuery并仅覆盖指定分区而非整张表?

仅覆盖BigQuery日期分区的Polars写入方案

需求:将Polars DataFrame写入按日期分区的BigQuery表,运行回填脚本时遍历日期范围,每日数据单独写入,但当前代码会覆盖整张表,需要实现仅覆盖对应日期的分区。

核心解决方案

  • 指定具体分区路径:BigQuery支持通过[表名]$YYYYMMDD的格式直接定位到单日期分区,在循环中根据当前处理的日期生成对应目标路径。
  • 复用写入配置但缩小作用范围:保留WRITE_TRUNCATE配置,当目标路径指向具体分区时,该配置只会截断并覆盖对应分区的数据,而非整张表。
  • 自动满足分区过滤要求:由于直接指定了分区路径,即使表开启了require_partition_filter=True,也无需额外添加过滤条件即可通过校验。

修改后的代码

import polars as pl
from google.cloud import bigquery
import io  # 补充导入io模块

# create period_range from internal util_package

for date in period_range:
    data =  "get some API data here per date"

    df = pl.read_csv(data).select(pl.col(pl.INT64))

    client = bigquery.Client()
    # 生成目标分区路径:表名拼接$YYYYMMDD格式的日期
    partition_date_str = date.strftime("%Y%m%d")
    destination_table = f"analytics.ads.vendor_name${partition_date_str}"
    
    with io.BytesIO() as stream:
        df.write_parquet(stream)
        stream.seek(0)
        job = client.load_table_from_file(
            file_obj=stream,
            destination=destination_table,
            project="mycompany_ads",
            location="EU",
            job_config=bigquery.LoadJobConfig(
                 source_format=bigquery.SourceFormat.PARQUET,
                 clustering_fields=["domain", "type", "placement"],
                 autodetect=True,
                 schema=None,
                 write_disposition=bigquery.WriteDisposition.WRITE_TRUNCATE,
           ),
      )
    job.result()  # 等待任务完成
    print(f"ETL finished for date: {date}")

关键修改说明

  1. 新增分区路径生成逻辑:将循环中的date格式化为YYYYMMDD字符串,拼接到表名后,明确指定写入目标为该日期的分区。
  2. 移除冗余的分区配置:写入具体分区时无需重复指定time_partitioning,因为表本身已配置好分区规则。
  3. 优化日志输出:添加日期标识,便于跟踪每日任务的完成情况。

内容的提问来源于stack exchange,提问作者Vega

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 02:15:56