如何在不下载的情况下获取远程归档文件的文件列表
远程归档文件文件名列表获取(Java实现方案)
核心思路
要避免下载完整大文件,关键是仅请求归档文件中存储目录/索引的字节范围,再用Apache Commons Compress解析这些片段。不同归档格式的目录存储位置差异很大,需针对性处理:
- ZIP:中央目录(Central Directory)位于文件尾部
- 7z:索引信息存储在文件尾部的
End Header中,可先请求尾部字节定位索引位置 - TAR:无集中目录,需遍历固定512字节的文件头,通过范围请求分段读取
- RAR:新版RAR5目录在尾部,旧版RAR目录分布在文件中;Apache Commons Compress仅支持RAR3及部分RAR5格式
基于Apache Commons Compress的实现步骤
1. 确认服务器支持HTTP范围请求
发送HEAD请求,检查响应头是否包含Accept-Ranges: bytes,不支持则无法使用此方案。
2. 分格式具体实现
ZIP格式
ZIP的中央目录末尾有固定22字节的End of Central Directory Record(EOCD),步骤如下:
import org.apache.commons.compress.archivers.zip.ZipArchiveInputStream; import org.apache.commons.compress.archivers.zip.ZipArchiveEntry; import java.net.HttpURLConnection; import java.io.InputStream; import java.net.URL; public class RemoteZipReader { public static void listZipEntries(String zipUrl) throws Exception { URL url = new URL(zipUrl); // 获取文件总大小 HttpURLConnection headConn = (HttpURLConnection) url.openConnection(); headConn.setRequestMethod("HEAD"); long fileSize = headConn.getContentLengthLong(); headConn.disconnect(); // 请求最后22字节的EOCD HttpURLConnection eocdConn = (HttpURLConnection) url.openConnection(); eocdConn.setRequestProperty("Range", "bytes=" + (fileSize - 22) + "-"); InputStream eocdIn = eocdConn.getInputStream(); byte[] eocdBytes = eocdIn.readAllBytes(); eocdConn.disconnect(); // 解析EOCD获取中央目录的偏移和大小 long cdOffset = ((eocdBytes[16] & 0xFF) << 24) | ((eocdBytes[17] & 0xFF) << 16) | ((eocdBytes[18] & 0xFF) << 8) | (eocdBytes[19] & 0xFF); long cdSize = ((eocdBytes[20] & 0xFF) << 24) | ((eocdBytes[21] & 0xFF) << 16) | ((eocdBytes[22] & 0xFF) << 8) | (eocdBytes[23] & 0xFF); // 请求中央目录数据并解析 HttpURLConnection cdConn = (HttpURLConnection) url.openConnection(); cdConn.setRequestProperty("Range", "bytes=" + cdOffset + "-" + (cdOffset + cdSize - 1)); InputStream cdIn = cdConn.getInputStream(); try (ZipArchiveInputStream zis = new ZipArchiveInputStream(cdIn)) { ZipArchiveEntry entry; while ((entry = zis.getNextZipEntry()) != null) { System.out.println(entry.getName()); } } cdConn.disconnect(); } }
7z格式
通过自定义SeekableByteChannel封装HTTP范围请求,让Apache Commons Compress可以像操作本地文件一样随机读取远程7z文件:
import org.apache.commons.compress.archivers.sevenz.SevenZFile; import java.nio.ByteBuffer; import java.nio.channels.SeekableByteChannel; import java.net.HttpURLConnection; import java.io.InputStream; import java.net.URL; class HttpSeekableByteChannel implements SeekableByteChannel { private final URL url; private long position = 0; private long fileSize; public HttpSeekableByteChannel(URL url) throws Exception { this.url = url; HttpURLConnection conn = (HttpURLConnection) url.openConnection(); conn.setRequestMethod("HEAD"); fileSize = conn.getContentLengthLong(); conn.disconnect(); } @Override public int read(ByteBuffer dst) throws Exception { long start = position; long end = Math.min(position + dst.remaining(), fileSize) - 1; if (start > end) return -1; HttpURLConnection conn = (HttpURLConnection) url.openConnection(); conn.setRequestProperty("Range", "bytes=" + start + "-" + end); try (InputStream in = conn.getInputStream()) { byte[] buffer = new byte[(int)(end - start + 1)]; int read = in.read(buffer); if (read > 0) { dst.put(buffer, 0, read); position += read; } return read; } finally { conn.disconnect(); } } @Override public long position() { return position; } @Override public SeekableByteChannel position(long newPosition) throws Exception { if (newPosition < 0 || newPosition > fileSize) throw new Exception("Invalid position"); this.position = newPosition; return this; } @Override public long size() { return fileSize; } @Override public int write(ByteBuffer src) { throw new UnsupportedOperationException(); } @Override public SeekableByteChannel truncate(long size) { throw new UnsupportedOperationException(); } @Override public boolean isOpen() { return true; } @Override public void close() {} } public class RemoteSevenZReader { public static void listSevenZEntries(String sevenzUrl) throws Exception { URL url = new URL(sevenzUrl); try (SeekableByteChannel channel = new HttpSeekableByteChannel(url); SevenZFile sevenZFile = new SevenZFile(channel)) { SevenZArchiveEntry entry; while ((entry = sevenZFile.getNextEntry()) != null) { System.out.println(entry.getName()); } } } }
TAR格式
TAR无集中目录,需逐个请求512字节的文件头并判断结束标记:
import org.apache.commons.compress.archivers.tar.TarArchiveEntry; import org.apache.commons.compress.archivers.tar.TarArchiveInputStream; import java.net.HttpURLConnection; import java.io.InputStream; import java.net.URL; import java.io.ByteArrayInputStream; public class RemoteTarReader { public static void listTarEntries(String tarUrl) throws Exception { URL url = new URL(tarUrl); HttpURLConnection headConn = (HttpURLConnection) url.openConnection(); headConn.setRequestMethod("HEAD"); long fileSize = headConn.getContentLengthLong(); headConn.disconnect(); long offset = 0; while (offset < fileSize) { // 请求当前文件头 HttpURLConnection conn = (HttpURLConnection) url.openConnection(); conn.setRequestProperty("Range", "bytes=" + offset + "-" + (offset + 511)); InputStream in = conn.getInputStream(); byte[] headerBytes = in.readAllBytes(); conn.disconnect(); // 判断是否为全零结束标记 boolean isEnd = true; for (byte b : headerBytes) { if (b != 0) { isEnd = false; break; } } if (isEnd) break; // 解析文件头并计算下一个偏移 try (TarArchiveInputStream tis = new TarArchiveInputStream(new ByteArrayInputStream(headerBytes))) { TarArchiveEntry entry = tis.getNextTarEntry(); if (entry != null) { System.out.println(entry.getName()); offset += 512 + ((entry.getSize() + 511) / 512) * 512; } else { offset += 512; } } } } }
类似remotezip的通用工具封装
上面的HttpSeekableByteChannel是核心,它实现了SeekableByteChannel接口,可直接被Apache Commons Compress的SevenZFile、ZipFile等类使用,模拟本地文件的随机访问,无需下载完整文件。你可以基于此封装支持多格式的远程归档阅读器。
注意事项
- 服务器兼容性:必须确保目标服务器支持HTTP范围请求,否则只能下载完整文件。
- 格式限制:Apache Commons Compress对RAR支持有限,如需完整RAR支持需使用第三方商业库。
- 边界处理:归档文件可能包含可变长度的注释(如ZIP的EOCD注释),解析时需考虑此类情况避免索引定位错误。
- 性能优化:批量请求字节范围可减少HTTP连接数,提升效率。
内容的提问来源于stack exchange,提问作者Alper M.
相关产品推荐
相关产品推荐

