You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

大栅格图像转DataFrame过慢,求高效方案及直接训练可行性

大尺寸栅格数组转DataFrame优化及直接训练Random Forest方案

问题背景

将大尺寸栅格图像数组转换为DataFrame用于Random Forest训练时,因图像尺寸过大导致DataFrame创建耗时过长。现有处理代码如下:

图像读取预处理代码

# Get the full paths of the specified images
images_3m = [os.path.join(folder_path, f) for f in specific_images]

# Define arrays to store 3m image data
arrays_3m = []

# Loop through each 3m image
for image in images_3m:
    print(image)
    src = rasterio.open(image)
    profile = src.profile
    print(profile)

    # Read 3m image data
    array = src.read(1)
    
    # Convert all images to float64 to handle LST3m and SAVI30m
    array = array.astype(np.float64)
    
    # Handle nodata values (replace with 0)
    array[array == -3.4028230607370965e+38] = 0
    #array[array == -1.797693e+308] = 0
    
    # Store 3m image data in arrays list
    arrays_3m.append(array)
    
    # Display the 3m image
    plt.imshow(array, cmap='magma')
    plt.colorbar()
    plt.show()  
    src.close()

DataFrame创建代码

from concurrent.futures import ThreadPoolExecutor

# Function to process a single image array into a list of pixel data
ny_3m = arrays_3m[0].shape[0]
nx_3m = arrays_3m[0].shape[1]

def process_array(i):
    values = []
    for y in range(ny_3m):
        for x in range(nx_3m):
            values.append([y, x, arrays_3m[i][y, x]])
    return np.array(values)

# Function to create DataFrame using threading
def create_dataframe_threading():
    with ThreadPoolExecutor() as executor:
        # Parallelize processing each array
        results = list(executor.map(process_array, range(len(arrays_3m))))
    
    values_3m = np.array(results)
    print(values_3m.shape)

    # Creating the DataFrame
    df_3m = pd.DataFrame(values_3m[:, :, 2].T, columns=['NDVI', 'MSAVI', 'SAVI', 'IPVI', 'DVI'])
    df_3m["y"] = values_3m[0][:, 0]
    df_3m["x"] = values_3m[0][:, 1]
    
    # Filter out nodata values
    df_3m = df_3m[df_3m["NDVI"] != 0]
    
    return df_3m

# Call the function to create the DataFrame
df_3m = create_dataframe_threading()
print(df_3m)

提问:是否有更快的栅格数组转DataFrame的方法?或者能否不创建DataFrame直接训练Random Forest模型?


一、更快的栅格数组转DataFrame方法

原代码的嵌套逐像素循环和线程处理效率极低,核心问题是未利用numpy的向量化运算优势。以下是优化方案:

优化后代码

import numpy as np
import pandas as pd
import rasterio
import os

# 读取并预处理图像
folder_path = "你的文件夹路径"
specific_images = ["NDVI.tif", "MSAVI.tif", "SAVI.tif", "IPVI.tif", "DVI.tif"]
arrays_3m = []

for image in specific_images:
    img_path = os.path.join(folder_path, image)
    with rasterio.open(img_path) as src:
        # 读取并转换数据类型
        array = src.read(1).astype(np.float64)
        # 处理nodata值
        array[array == -3.4028230607370965e+38] = 0
        arrays_3m.append(array)

# 生成坐标网格(向量化操作,避免循环)
ny_3m, nx_3m = arrays_3m[0].shape
y_coords, x_coords = np.mgrid[0:ny_3m, 0:nx_3m]

# 展平所有波段数组并合并为特征矩阵
features = np.column_stack([arr.flatten() for arr in arrays_3m])
# 展平坐标并合并
coords = np.column_stack([y_coords.flatten(), x_coords.flatten()])

# 创建DataFrame
df_3m = pd.DataFrame(features, columns=['NDVI', 'MSAVI', 'SAVI', 'IPVI', 'DVI'])
df_3m[['y', 'x']] = coords

# 过滤nodata值
df_3m = df_3m[df_3m['NDVI'] != 0]

核心优化点

  • 用np.mgrid直接生成坐标网格,替代逐像素循环,速度提升数量级
  • 利用数组flatten()和column_stack完成数据合并,完全依赖numpy底层优化的向量化运算,避免Python循环开销
  • 移除不必要的线程处理:CPU密集型数组操作中,线程切换开销大于并行收益,numpy已做底层优化

二、不创建DataFrame,直接训练Random Forest模型

Sklearn的Random Forest模型原生支持numpy数组输入,无需通过DataFrame中转,可直接用特征矩阵和标签数组训练,大幅节省内存和时间。

直接训练代码示例

假设存在标签栅格(如LST3m)作为训练目标:

from sklearn.ensemble import RandomForestRegressor

# 读取并预处理标签栅格
with rasterio.open(os.path.join(folder_path, "LST3m.tif")) as src:
    label_array = src.read(1).astype(np.float64)
    label_array[label_array == -3.4028230607370965e+38] = 0

# 生成过滤掩码(同时过滤特征和标签的nodata值)
mask = (arrays_3m[0].flatten() != 0) & (label_array.flatten() != 0)

# 生成训练用的特征矩阵X和标签数组y
X = np.column_stack([arr.flatten() for arr in arrays_3m])[mask]
y = label_array.flatten()[mask]

# 训练Random Forest模型(n_jobs=-1利用所有CPU核心)
rf_model = RandomForestRegressor(n_estimators=100, n_jobs=-1, random_state=42)
rf_model.fit(X, y)

优势

  • 完全跳过DataFrame创建步骤,消除额外内存占用,适合超大栅格场景
  • 直接用numpy数组训练,与sklearn模型无缝兼容,效率更高
  • n_jobs=-1开启并行训练,进一步缩短模型训练时间

内容的提问来源于stack exchange,提问作者Budin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 10:09:51