如何从JAR包内读取Avro Parquet资源文件且无需生成临时文件
解决方案
问题根因
报错org.apache.hadoop.fs.UnsupportedFileSystemException: No FileSystem for scheme "jar"是因为Hadoop FileSystem默认未实现jar协议的处理器,无法直接解析JAR包内的资源路径。而ParquetReader读取文件需要支持随机定位(seek操作)来解析文件尾部的元数据,所以不能直接传入普通InputStream。
实现思路
自定义实现Parquet的InputFile接口,将JAR包内读取到的资源流全量加载到内存字节数组中,封装为支持seek操作的输入源,全程无需写出到临时文件。
完整可运行代码
1. 自定义ByteArrayInputFile实现
import org.apache.parquet.io.InputFile; import org.apache.parquet.io.SeekableInputStream; import java.io.IOException; import java.nio.ByteBuffer; public class ByteArrayInputFile implements InputFile { private final byte[] data; public ByteArrayInputFile(byte[] data) { this.data = data; } @Override public long getLength() { return data.length; } @Override public SeekableInputStream newStream() { return new SeekableInputStream() { private final ByteBuffer buf = ByteBuffer.wrap(data); @Override public long getPos() { return buf.position(); } @Override public void seek(long l) throws IOException { if (l < 0 || l >= data.length) { throw new IOException("无效的定位位置: " + l); } buf.position((int) l); } @Override public int read() { if (!buf.hasRemaining()) { return -1; } return buf.get() & 0xFF; } @Override public int read(byte[] b, int off, int len) { if (!buf.hasRemaining()) { return -1; } int readLen = Math.min(len, buf.remaining()); buf.get(b, off, readLen); return readLen; } @Override public void close() {} }; } }
2. 改造原有读取逻辑
import org.apache.avro.generic.GenericRecord; import org.apache.hadoop.conf.Configuration; import org.apache.parquet.avro.AvroParquetReader; import org.apache.parquet.hadoop.ParquetReader; import java.io.ByteArrayOutputStream; import java.io.IOException; import java.io.InputStream; // 你的业务逻辑部分 ClassLoader classLoader = Thread.currentThread().getContextClassLoader(); try (InputStream resourceStream = classLoader.getResourceAsStream(pattern_id); ByteArrayOutputStream baos = new ByteArrayOutputStream()) { if (resourceStream == null) { throw new IOException("未找到指定资源: " + pattern_id); } // 把Parquet资源全量读入字节数组 byte[] buffer = new byte[4096]; int bytesRead; while ((bytesRead = resourceStream.read(buffer)) != -1) { baos.write(buffer, 0, bytesRead); } byte[] parquetData = baos.toByteArray(); // 构造ParquetReader try (ParquetReader<GenericRecord> reader = AvroParquetReader.<GenericRecord>builder(new ByteArrayInputFile(parquetData)) .disableCompatibility() .withConf(new Configuration()) .build()) { patternsFound.add(pattern_id); GenericRecord record; while ((record = reader.read()) != null) { // 你的原有业务处理逻辑 } } } catch (IOException e) { e.printStackTrace(); }
注意事项
- 本方案将Parquet文件全量加载到内存,仅适合内嵌在JAR中的小体积Parquet资源使用,若文件体积超过512M建议还是写出临时文件避免OOM
- 无需新增额外依赖,兼容你原有的Parquet、Avro、Hadoop依赖版本
内容的提问来源于stack exchange,提问作者user3112960
相关产品推荐
相关产品推荐

