You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中遍历Azure Blob存储的Parquet文件并合并?

解决方法

步骤1:安装并加载必要的R包

需要用到AzureStor访问Blob存储,arrow读写Parquet文件:

install.packages(c("AzureStor", "arrow"))
library(AzureStor)
library(arrow)

步骤2:连接到Azure Blob存储容器

替换为你的存储账户信息和容器名称:

# 存储账户基础信息
account_name <- "myacct"
account_key <- "wRGSJ"

# 创建存储端点并获取容器对象
endp <- storage_endpoint(paste0("https://", account_name, ".blob.core.windows.net"), key = account_key)
cont <- storage_container(endp, "retail") # 这里替换为你的容器名称

步骤3:筛选容器内所有.parquet文件

列出容器中后缀为.parquet的所有文件:

parquet_files <- list_storage_files(cont, pattern = "\\.parquet$")

步骤4:批量读取并合并文件

如果数据量较小,可直接用循环合并;数据量大的话推荐用dplyr提升效率:

基础循环方式

combined_df <- data.frame()

for (file in parquet_files) {
  file_url <- storage_url(cont, file$name)
  raw_data <- download_from_url(file_url, key = account_key, dest = NULL)
  parq_df <- read_parquet(raw_data)
  
  combined_df <- rbind(combined_df, parq_df)
}

高效合并方式(推荐)

library(dplyr)

combined_df <- lapply(parquet_files, function(file) {
  file_url <- storage_url(cont, file$name)
  raw_data <- download_from_url(file_url, key = account_key, dest = NULL)
  read_parquet(raw_data)
}) %>% bind_rows()

步骤5:写入合并后的单个Parquet文件

write_parquet(combined_df, "combined_retail.parquet")

注意事项

  • 确保你的存储密钥拥有容器的读取权限
  • 若文件结构不一致(比如列名/数据类型不同),合并前需先统一数据结构
  • 超大规模数据可考虑分批次处理,避免内存溢出

内容的提问来源于stack exchange,提问作者Surely

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 13:04:57