You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过Java程序化将TSV/CSV转换为带Snappy压缩的ORC文件?

Can I Convert TSV/CSV to Snappy-Compressed ORC via Java?

Absolutely! You can absolutely convert TSV/CSV files to Snappy-compressed ORC files using Java—this is a common task in big data pipelines, and there are robust libraries to make this straightforward. Below is a step-by-step guide with working code to get you started.

Prerequisites

You'll need two key libraries to pull this off:

  • Apache ORC: The core library for creating and manipulating ORC files.
  • Apache Commons CSV: A reliable tool for parsing TSV/CSV files (avoids the hassle of manual string splitting and handling edge cases like quoted fields or escaped delimiters).

Maven Dependencies

Add these to your pom.xml (use the latest stable versions available for best compatibility):

<dependencies>
    <!-- Apache ORC Core Library -->
    <dependency>
        <groupId>org.apache.orc</groupId>
        <artifactId>orc-core</artifactId>
        <version>1.8.2</version>
    </dependency>
    <!-- Apache Commons CSV for TSV/CSV Parsing -->
    <dependency>
        <groupId>org.apache.commons</groupId>
        <artifactId>commons-csv</artifactId>
        <version>1.10.0</version>
    </dependency>
</dependencies>

Full Implementation Code

This example assumes your CSV has a header row with fields id (int), name (string), age (int), and score (double). Adjust the schema and parsing logic to match your actual TSV/CSV structure.

import org.apache.commons.csv.CSVFormat;
import org.apache.commons.csv.CSVParser;
import org.apache.commons.csv.CSVRecord;
import org.apache.orc.OrcFile;
import org.apache.orc.TypeDescription;
import org.apache.orc.Writer;

import java.io.FileReader;
import java.io.IOException;
import java.util.ArrayList;
import java.util.List;

public class CsvToSnappyOrcConverter {

    public static void main(String[] args) throws IOException {
        // 1. Define the ORC schema (must match your input file's fields)
        TypeDescription orcSchema = TypeDescription.fromString(
            "struct<id:int,name:string,age:int,score:double>"
        );

        // 2. Configure ORC Writer with Snappy compression enabled
        OrcFile.WriterOptions writerOpts = OrcFile.writerOptions(null)
            .setSchema(orcSchema)
            .compress(OrcFile.CompressionKind.SNAPPY); // This line enables Snappy compression

        // 3. Use try-with-resources to auto-close resources and avoid leaks
        try (Writer orcWriter = OrcFile.createWriter(
                new org.apache.hadoop.fs.Path("output.snappy.orc"), 
                writerOpts
            );
            CSVParser csvParser = new CSVParser(
                new FileReader("input.csv"), 
                CSVFormat.DEFAULT.withHeader() // Swap to CSVFormat.TDF for TSV files
            )) {

            // 4. Parse each CSV row and write it to the ORC file
            for (CSVRecord csvRow : csvParser) {
                List<Object> orcRow = new ArrayList<>();
                orcRow.add(Integer.parseInt(csvRow.get("id")));
                orcRow.add(csvRow.get("name"));
                orcRow.add(Integer.parseInt(csvRow.get("age")));
                orcRow.add(Double.parseDouble(csvRow.get("score")));

                orcWriter.addRow(orcRow);
            }

            System.out.println("Conversion complete! Snappy-compressed ORC file generated.");
        }
    }
}

Key Details to Note

  • Schema Definition: ORC is a columnar storage format, so you must define a schema that matches your input file's fields. ORC supports a wide range of types (including arrays, maps, and nested structs) for complex data structures.
  • Snappy Compression: The line .compress(OrcFile.CompressionKind.SNAPPY) is critical—it enables Snappy compression for the output ORC file. ORC also supports ZLIB, LZ4, and ZSTD if you need different compression tradeoffs.
  • TSV Support: To parse TSV files instead of CSV, replace CSVFormat.DEFAULT.withHeader() with CSVFormat.TDF.withHeader() (TDF stands for Tab-Delimited Format).
  • Performance Tuning: For large files, you can adjust ORC writer parameters like setStripeSize (controls the size of each stripe in the ORC file) or setBufferSize to optimize memory usage and write speed.
  • Resource Management: Using try-with-resources ensures that the ORC writer and CSV parser are closed properly, preventing resource leaks.

Troubleshooting Tips

  • If you encounter Hadoop-related errors (ORC depends on Hadoop APIs), ensure your project includes the necessary Hadoop dependencies (e.g., hadoop-common) or use the shaded ORC jar (orc-shaded) to avoid dependency conflicts.
  • For very large files, consider processing the input in batches to avoid loading the entire file into memory at once.

内容的提问来源于stack exchange,提问作者Amit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:43:14