You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将新增CSV文件的数据加载至已创建的Delta表?

How to Load New CSV Data into an Existing Delta Table

Hey there! Since you’ve already set up your Delta table using a CSV file with that initial SQL code, adding new CSV data is super straightforward—you can leverage the same COPY INTO command you used before, with just a few key notes to make it smooth.

Step-by-Step Implementation

First, let’s recap the core approach: COPY INTO is built exactly for this incremental load scenario, and it’s idempotent (meaning running it multiple times won’t re-load the same data, so you don’t have to stress about duplicates).

Here’s the code to load your new CSV file:

%sql
-- Optional: Keep this if you need to handle new columns in the incoming CSV
SET spark.databricks.delta.schema.autoMerge.enabled = true;

-- Load the new CSV data into your existing Delta table
COPY INTO catalog.schema.tablename
  FROM (SELECT * FROM 's3://bucket/your_new_data.csv') 
    FILEFORMAT = CSV 
    FORMAT_OPTIONS ('mergeSchema' = 'true', 'header' = 'true')

Key Tips & Variations

  • Load multiple files at once: If your new CSVs are stored in the same S3 prefix (e.g., s3://bucket/csv_incrementals/), you can point COPY INTO directly to that folder instead of individual files. It will automatically load all unprocessed CSV files in that path:
    COPY INTO catalog.schema.tablename
      FROM (SELECT * FROM 's3://bucket/csv_incrementals/') 
        FILEFORMAT = CSV 
        FORMAT_OPTIONS ('mergeSchema' = 'true', 'header' = 'true')
    
  • Transform data before loading: If you need to clean up data, adjust data types, or rename columns before loading, modify the SELECT clause in the source query:
    COPY INTO catalog.schema.tablename
      FROM (
        SELECT 
          customer_id,
          cast(order_amount as decimal(10,2)) as order_amount, -- Convert data type
          new_user_flag -- Automatically adds this new column if it doesn't exist
        FROM 's3://bucket/your_new_data.csv'
      ) 
        FILEFORMAT = CSV 
        FORMAT_OPTIONS ('mergeSchema' = 'true', 'header' = 'true')
    
  • Schema compatibility: The mergeSchema=true flag ensures that any new columns in your incoming CSV are automatically added to the Delta table. If your new CSV has the same columns as the existing table, this flag still works safely (it just won’t make any schema changes).

内容的提问来源于stack exchange,提问作者patdev

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 18:10:31