You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Hive ORC文件格式:压缩数据如何转换为可读格式?

How Hive Converts Compressed ORC Data to Readable Format During SELECT * Queries

Great question! Let’s walk through exactly how Hive turns those compressed, unreadable ORC files in HDFS into the clean, readable results you get from a SELECT * query.

First, a quick primer: ORC (Optimized Row Columnar) is a columnar storage format with compression baked into its core design. When you create an ORC table in Hive, it automatically compresses data at the column level (the default codec is ZLIB, though you can configure alternatives like Snappy or LZO). That’s why the files in HDFS look like binary gibberish if you try to read them directly—they’re compressed, columnar binary data, not plain text.

Here’s the step-by-step breakdown of what happens when you run your query:

  • Step 1: Query Planning & Data Locality
    Hive’s execution engine (Tez, MapReduce, or Spark, depending on your setup) first parses your SELECT * query, identifies the relevant ORC files in HDFS, and schedules tasks to run on the nodes where those files are stored. This "data locality" minimizes unnecessary data transfer across the cluster, speeding up the process.

  • Step 2: ORC File Parsing & Decompression
    Each task uses Hive’s built-in ORC File Reader to process the compressed files:

    1. It first reads the ORC file’s footer and metadata, which holds critical details: column data types, the compression codec used, the location of each "stripe" (ORC groups rows into stripes for efficiency), and data statistics.
    2. For each stripe, the reader uses the specified codec to decompress the column data. Since ORC compresses columns individually, even though SELECT * pulls all columns, it only decompresses the ones present in your table—no wasted effort on unused data.
  • Step 3: Columnar-to-Row Conversion
    ORC stores data in columnar format (all values for a single column are grouped together), but SQL queries expect row-based results. The ORC reader takes the decompressed column data and reconstructs full rows by matching corresponding values across columns. For example, if your table has user_id, username, and email, it pairs the first user_id with the first username and email to make a complete row, and repeats this for all entries.

  • Step 4: Formatting & Delivering Results
    Once rows are reconstructed, the execution engine sends them back to Hive’s driver, which formats them into a human-readable output—like tab-separated values in the CLI, or a structured table view if you’re using a UI like Hue.

A bonus efficiency note: Even with decompression and format conversion, ORC’s design keeps this process fast. Column-level compression reduces file sizes (so less data is read from HDFS), and modern codecs like Snappy are optimized for quick decompression—so the overhead is minimal compared to reading uncompressed data.

内容的提问来源于stack exchange,提问作者Aditya Singh Bhadoria

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:58:57