Hive ORC文件格式:压缩数据如何转换为可读格式?
SELECT * Queries Great question! Let’s walk through exactly how Hive turns those compressed, unreadable ORC files in HDFS into the clean, readable results you get from a SELECT * query.
First, a quick primer: ORC (Optimized Row Columnar) is a columnar storage format with compression baked into its core design. When you create an ORC table in Hive, it automatically compresses data at the column level (the default codec is ZLIB, though you can configure alternatives like Snappy or LZO). That’s why the files in HDFS look like binary gibberish if you try to read them directly—they’re compressed, columnar binary data, not plain text.
Here’s the step-by-step breakdown of what happens when you run your query:
Step 1: Query Planning & Data Locality
Hive’s execution engine (Tez, MapReduce, or Spark, depending on your setup) first parses yourSELECT *query, identifies the relevant ORC files in HDFS, and schedules tasks to run on the nodes where those files are stored. This "data locality" minimizes unnecessary data transfer across the cluster, speeding up the process.Step 2: ORC File Parsing & Decompression
Each task uses Hive’s built-in ORC File Reader to process the compressed files:- It first reads the ORC file’s footer and metadata, which holds critical details: column data types, the compression codec used, the location of each "stripe" (ORC groups rows into stripes for efficiency), and data statistics.
- For each stripe, the reader uses the specified codec to decompress the column data. Since ORC compresses columns individually, even though
SELECT *pulls all columns, it only decompresses the ones present in your table—no wasted effort on unused data.
Step 3: Columnar-to-Row Conversion
ORC stores data in columnar format (all values for a single column are grouped together), but SQL queries expect row-based results. The ORC reader takes the decompressed column data and reconstructs full rows by matching corresponding values across columns. For example, if your table hasuser_id,username, andemail, it pairs the firstuser_idwith the firstusernameandemailto make a complete row, and repeats this for all entries.Step 4: Formatting & Delivering Results
Once rows are reconstructed, the execution engine sends them back to Hive’s driver, which formats them into a human-readable output—like tab-separated values in the CLI, or a structured table view if you’re using a UI like Hue.
A bonus efficiency note: Even with decompression and format conversion, ORC’s design keeps this process fast. Column-level compression reduces file sizes (so less data is read from HDFS), and modern codecs like Snappy are optimized for quick decompression—so the overhead is minimal compared to reading uncompressed data.
内容的提问来源于stack exchange,提问作者Aditya Singh Bhadoria

