You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用rasterio分块将大型Dask数组上传至S3?

解决Dask数组分块写入S3 COG仅单块生效的问题

你的代码核心错误是**Window.from_slices的切片范围定义错误**:你传入的rows=(y0, y_chunk)意味着结束行是y_chunk而非y0 + y_chunk,cols参数同理。这会导致所有块都被写入到同一个左上角区域,最终只有最后一块的数据能被保留。

修正方案:

1. 基础修正:修复窗口范围

直接修改窗口的结束位置计算逻辑,确保每个块写入到正确的地理范围:

import rasterio
from rasterio.windows import Window

s3_path = "/vsis3/some_bucket/some_file.tif"

# 获取Dask数组的块大小,显式写入COG配置
chunk_y, chunk_x = tif_clipped.chunksizes["y"][0], tif_clipped.chunksizes["x"][0]

profile = {
    "driver": "COG",
    "height": tif_clipped.sizes["y"],
    "width": tif_clipped.sizes["x"],
    "count": 4,
    "tiled": True,
    # 显式设置块大小,与Dask块对齐,避免额外重采样
    "blockxsize": chunk_x,
    "blockysize": chunk_y,
    "compress": "lzw",
    "dtype": tif_clipped.dtype,
    "crs": tif_clipped.rio.crs,
    "transform": tif_clipped.rio.transform(),
}

with rasterio.Env(CPL_VSIL_USE_TEMP_FILE_FOR_RANDOM_WRITE="YES"):
    with rasterio.open(s3_path, "w", **profile) as dst:
        chunks_y = tif_clipped.chunksizes["y"]
        chunks_x = tif_clipped.chunksizes["x"]
        total = len(chunks_x) * len(chunks_y)
        counter = 0

        for y_off, y_chunk in enumerate(chunks_y):
            y0 = sum(chunks_y[:y_off])
            y_end = y0 + y_chunk  # 修正:结束行=起始行+块大小
            for x_off, x_chunk in enumerate(chunks_x):
                counter += 1
                print(f"Uploading chunk {counter}/{total}")
                x0 = sum(chunks_x[:x_off])
                x_end = x0 + x_chunk  # 修正:结束列=起始列+块大小

                window = Window.from_slices(rows=(y0, y_end), cols=(x0, x_end))

                data_chunk = tif_clipped.isel(
                    y=slice(y0, y_end), x=slice(x0, x_end)
                ).compute()

                dst.write(data_chunk.values, window=window)

2. 优化方案:利用xarray内置方法简化逻辑

使用rio.iter_chunks()直接获取每个块的窗口和对应数据,避免手动计算起始位置,减少出错概率:

import rasterio

s3_path = "/vsis3/some_bucket/some_file.tif"

chunk_y, chunk_x = tif_clipped.chunksizes["y"][0], tif_clipped.chunksizes["x"][0]
profile = {
    "driver": "COG",
    "height": tif_clipped.sizes["y"],
    "width": tif_clipped.sizes["x"],
    "count": 4,
    "tiled": True,
    "blockxsize": chunk_x,
    "blockysize": chunk_y,
    "compress": "lzw",
    "dtype": tif_clipped.dtype,
    "crs": tif_clipped.rio.crs,
    "transform": tif_clipped.rio.transform(),
}

with rasterio.Env(CPL_VSIL_USE_TEMP_FILE_FOR_RANDOM_WRITE="YES"):
    with rasterio.open(s3_path, "w", **profile) as dst:
        total_chunks = len(list(tif_clipped.rio.iter_chunks()))
        for idx, (window, block) in enumerate(tif_clipped.rio.iter_chunks()):
            print(f"Uploading chunk {idx+1}/{total_chunks}")
            # 计算当前块并写入对应窗口
            dst.write(block.compute().values, window=window)

关键说明:

  • 显式设置blockxsize和blockysize可以让COG的内部块与Dask块完全对齐,避免写入时的额外重采样操作。
  • CPL_VSIL_USE_TEMP_FILE_FOR_RANDOM_WRITE="YES"的配置是必要的,确保rasterio能在S3对象存储上支持随机写入操作。

内容的提问来源于stack exchange,提问作者ultrapoci

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 17:07:09