You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Drill读取Azure存储中Parquet文件的IO性能问题及优化咨询

Root Cause Analysis

Based on your detailed benchmarks and observations, here are the key reasons behind the poor Parquet read performance from Azure storage in Drill:

  1. Mismatched IO Pattern with Azure Storage
    Parquet is a columnar format, and Drill's default Parquet reader tends to use small, random IO requests to fetch individual column blocks. Azure Blob/Files (especially SMB-mounted Premium Files) perform poorly with random IO compared to local SSDs—each small request adds network/protocol overhead that accumulates into drastically lower throughput. CSV, by contrast, uses sequential IO which aligns perfectly with Azure storage's strengths, hence the normal performance.

  2. Excessive Unnecessary Data Transfer
    The 2.5GB outgoing traffic for a 300MB Parquet file indicates Drill is either:

    • Failing to leverage Parquet's built-in statistics (min/max values) to skip irrelevant column blocks, leading to full file scans.
    • Repeatedly fetching the same data blocks due to inefficient caching or metadata handling.
    • Generating redundant API calls to Azure Blob, which amplifies network overhead.
  3. SMB Protocol Overhead for Azure Files
    While your Premium Files benchmark shows strong performance with large blocks, the SMB protocol adds inherent latency for small, frequent IO requests. Drill's Parquet reader's default behavior floods the SMB mount with these small requests, negating the underlying storage's capabilities.

  4. Lack of Azure-Specific Optimizations
    Drill's out-of-the-box Parquet configuration isn't tuned for Azure storage's unique characteristics (like optimal block sizes, batch read limits, or Blob API constraints). Local storage benefits from lower latency and more efficient IO paths, which is why Parquet performs normally there.


Actionable Optimization Solutions

1. Tune Drill's Parquet Reader Configuration

Modify drill-override.conf to align with Azure's performance sweet spots:

drill.exec: {
  parquet: {
    # Increase block size to match Azure's optimal large-block IO (64MB)
    reader.block.size: 67108864,
    # Disable small file merging if your Parquet files are already large (>64MB)
    reader.merge_small_files: false
  },
  planner: {
    # Ensure predicate pushdown and column pruning are enabled to avoid full scans
    enable_parquet_predicate_pushdown: true,
    enable_column_pruning: true
  }
}

2. Optimize Azure Blob Storage Access

Skip SMB mounts for Blob storage and use Drill's native Azure Blob plugin with batch read settings. Update your Blob storage plugin configuration:

{
  "type": "file",
  "connection": "az://your-container/",
  "config": {
    "fs.azure.account.key.your-storage-account.blob.core.windows.net": "your-access-key"
  },
  "formats": {
    "parquet": {
      "type": "parquet",
      "block-size": 67108864,
      "read-ahead-size": 134217728
    }
  }
}

If possible, switch to Azure Data Lake Storage Gen2—it supports hierarchical namespaces and optimized object storage access that works better with Drill's Parquet reader.

3. Optimize Azure Files (SMB) Mount

Adjust your SMB mount parameters to reduce small IO overhead:

sudo mount -t cifs //your-storage-account.file.core.windows.net/your-share /data/dfs \
-o vers=3.0,username=your-storage-account,password=your-access-key, \
uid=1000,gid=1000,file_mode=0777,dir_mode=0777, \
cache=strict,rsize=1048576,wsize=1048576,mtu=1500

The rsize/wsize settings (1MB) align with Azure Files' optimal block size, reducing the number of IO requests.

4. Drill Runtime Tuning

  • Increase Heap Allocation: Give Drill more memory to cache Parquet metadata and data blocks, reducing repeated IO. Update drill-env.sh:
    export DRILL_HEAP="24G" # Allocate 24GB of your VM's 32GB memory
    
  • Adjust Thread Pools: Match thread counts to your VM's vCPU resources (4-8 threads) to avoid context-switching overhead:
    drill.exec: {
      threads: {
        max: 8,
        min: 4
      }
    }
    

5. Validate Parquet File Structure

  • Use parquet-tools to check for efficient compression (Snappy/ZSTD are ideal) and intact metadata:
    parquet-tools meta your-parquet-file.parquet
    
  • Merge small Parquet files into larger ones (>64MB) to reduce the number of IO requests Drill needs to make.

内容的提问来源于stack exchange,提问作者rudolfdobias

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:57:52