You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从JAR包内读取Avro Parquet资源文件且无需生成临时文件

解决方案

问题根因

报错org.apache.hadoop.fs.UnsupportedFileSystemException: No FileSystem for scheme "jar"是因为Hadoop FileSystem默认未实现jar协议的处理器,无法直接解析JAR包内的资源路径。而ParquetReader读取文件需要支持随机定位(seek操作)来解析文件尾部的元数据,所以不能直接传入普通InputStream。

实现思路

自定义实现Parquet的InputFile接口,将JAR包内读取到的资源流全量加载到内存字节数组中,封装为支持seek操作的输入源,全程无需写出到临时文件。

完整可运行代码

1. 自定义ByteArrayInputFile实现

import org.apache.parquet.io.InputFile;
import org.apache.parquet.io.SeekableInputStream;
import java.io.IOException;
import java.nio.ByteBuffer;

public class ByteArrayInputFile implements InputFile {
    private final byte[] data;

    public ByteArrayInputFile(byte[] data) {
        this.data = data;
    }

    @Override
    public long getLength() {
        return data.length;
    }

    @Override
    public SeekableInputStream newStream() {
        return new SeekableInputStream() {
            private final ByteBuffer buf = ByteBuffer.wrap(data);

            @Override
            public long getPos() {
                return buf.position();
            }

            @Override
            public void seek(long l) throws IOException {
                if (l < 0 || l >= data.length) {
                    throw new IOException("无效的定位位置: " + l);
                }
                buf.position((int) l);
            }

            @Override
            public int read() {
                if (!buf.hasRemaining()) {
                    return -1;
                }
                return buf.get() & 0xFF;
            }

            @Override
            public int read(byte[] b, int off, int len) {
                if (!buf.hasRemaining()) {
                    return -1;
                }
                int readLen = Math.min(len, buf.remaining());
                buf.get(b, off, readLen);
                return readLen;
            }

            @Override
            public void close() {}
        };
    }
}

2. 改造原有读取逻辑

import org.apache.avro.generic.GenericRecord;
import org.apache.hadoop.conf.Configuration;
import org.apache.parquet.avro.AvroParquetReader;
import org.apache.parquet.hadoop.ParquetReader;
import java.io.ByteArrayOutputStream;
import java.io.IOException;
import java.io.InputStream;

// 你的业务逻辑部分
ClassLoader classLoader = Thread.currentThread().getContextClassLoader();
try (InputStream resourceStream = classLoader.getResourceAsStream(pattern_id);
     ByteArrayOutputStream baos = new ByteArrayOutputStream()) {

    if (resourceStream == null) {
        throw new IOException("未找到指定资源: " + pattern_id);
    }

    // 把Parquet资源全量读入字节数组
    byte[] buffer = new byte[4096];
    int bytesRead;
    while ((bytesRead = resourceStream.read(buffer)) != -1) {
        baos.write(buffer, 0, bytesRead);
    }
    byte[] parquetData = baos.toByteArray();

    // 构造ParquetReader
    try (ParquetReader<GenericRecord> reader = AvroParquetReader.<GenericRecord>builder(new ByteArrayInputFile(parquetData))
            .disableCompatibility()
            .withConf(new Configuration())
            .build()) {

        patternsFound.add(pattern_id);
        GenericRecord record;
        while ((record = reader.read()) != null) {
            // 你的原有业务处理逻辑
        }
    }
} catch (IOException e) {
    e.printStackTrace();
}

注意事项

  • 本方案将Parquet文件全量加载到内存,仅适合内嵌在JAR中的小体积Parquet资源使用,若文件体积超过512M建议还是写出临时文件避免OOM
  • 无需新增额外依赖,兼容你原有的Parquet、Avro、Hadoop依赖版本

内容的提问来源于stack exchange,提问作者user3112960

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 06:36:03