You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在不下载的情况下获取远程归档文件的文件列表

远程归档文件文件名列表获取(Java实现方案)

核心思路

要避免下载完整大文件,关键是仅请求归档文件中存储目录/索引的字节范围,再用Apache Commons Compress解析这些片段。不同归档格式的目录存储位置差异很大,需针对性处理:

  • ZIP:中央目录(Central Directory)位于文件尾部
  • 7z:索引信息存储在文件尾部的End Header中,可先请求尾部字节定位索引位置
  • TAR:无集中目录,需遍历固定512字节的文件头,通过范围请求分段读取
  • RAR:新版RAR5目录在尾部,旧版RAR目录分布在文件中;Apache Commons Compress仅支持RAR3及部分RAR5格式

基于Apache Commons Compress的实现步骤

1. 确认服务器支持HTTP范围请求

发送HEAD请求,检查响应头是否包含Accept-Ranges: bytes,不支持则无法使用此方案。

2. 分格式具体实现

ZIP格式

ZIP的中央目录末尾有固定22字节的End of Central Directory Record(EOCD),步骤如下:

import org.apache.commons.compress.archivers.zip.ZipArchiveInputStream;
import org.apache.commons.compress.archivers.zip.ZipArchiveEntry;
import java.net.HttpURLConnection;
import java.io.InputStream;
import java.net.URL;

public class RemoteZipReader {
    public static void listZipEntries(String zipUrl) throws Exception {
        URL url = new URL(zipUrl);
        // 获取文件总大小
        HttpURLConnection headConn = (HttpURLConnection) url.openConnection();
        headConn.setRequestMethod("HEAD");
        long fileSize = headConn.getContentLengthLong();
        headConn.disconnect();

        // 请求最后22字节的EOCD
        HttpURLConnection eocdConn = (HttpURLConnection) url.openConnection();
        eocdConn.setRequestProperty("Range", "bytes=" + (fileSize - 22) + "-");
        InputStream eocdIn = eocdConn.getInputStream();
        byte[] eocdBytes = eocdIn.readAllBytes();
        eocdConn.disconnect();

        // 解析EOCD获取中央目录的偏移和大小
        long cdOffset = ((eocdBytes[16] & 0xFF) << 24) | ((eocdBytes[17] & 0xFF) << 16) |
                        ((eocdBytes[18] & 0xFF) << 8) | (eocdBytes[19] & 0xFF);
        long cdSize = ((eocdBytes[20] & 0xFF) << 24) | ((eocdBytes[21] & 0xFF) << 16) |
                      ((eocdBytes[22] & 0xFF) << 8) | (eocdBytes[23] & 0xFF);

        // 请求中央目录数据并解析
        HttpURLConnection cdConn = (HttpURLConnection) url.openConnection();
        cdConn.setRequestProperty("Range", "bytes=" + cdOffset + "-" + (cdOffset + cdSize - 1));
        InputStream cdIn = cdConn.getInputStream();

        try (ZipArchiveInputStream zis = new ZipArchiveInputStream(cdIn)) {
            ZipArchiveEntry entry;
            while ((entry = zis.getNextZipEntry()) != null) {
                System.out.println(entry.getName());
            }
        }
        cdConn.disconnect();
    }
}

7z格式

通过自定义SeekableByteChannel封装HTTP范围请求,让Apache Commons Compress可以像操作本地文件一样随机读取远程7z文件:

import org.apache.commons.compress.archivers.sevenz.SevenZFile;
import java.nio.ByteBuffer;
import java.nio.channels.SeekableByteChannel;
import java.net.HttpURLConnection;
import java.io.InputStream;
import java.net.URL;

class HttpSeekableByteChannel implements SeekableByteChannel {
    private final URL url;
    private long position = 0;
    private long fileSize;

    public HttpSeekableByteChannel(URL url) throws Exception {
        this.url = url;
        HttpURLConnection conn = (HttpURLConnection) url.openConnection();
        conn.setRequestMethod("HEAD");
        fileSize = conn.getContentLengthLong();
        conn.disconnect();
    }

    @Override
    public int read(ByteBuffer dst) throws Exception {
        long start = position;
        long end = Math.min(position + dst.remaining(), fileSize) - 1;
        if (start > end) return -1;

        HttpURLConnection conn = (HttpURLConnection) url.openConnection();
        conn.setRequestProperty("Range", "bytes=" + start + "-" + end);
        try (InputStream in = conn.getInputStream()) {
            byte[] buffer = new byte[(int)(end - start + 1)];
            int read = in.read(buffer);
            if (read > 0) {
                dst.put(buffer, 0, read);
                position += read;
            }
            return read;
        } finally {
            conn.disconnect();
        }
    }

    @Override public long position() { return position; }
    @Override public SeekableByteChannel position(long newPosition) throws Exception {
        if (newPosition < 0 || newPosition > fileSize) throw new Exception("Invalid position");
        this.position = newPosition;
        return this;
    }
    @Override public long size() { return fileSize; }
    @Override public int write(ByteBuffer src) { throw new UnsupportedOperationException(); }
    @Override public SeekableByteChannel truncate(long size) { throw new UnsupportedOperationException(); }
    @Override public boolean isOpen() { return true; }
    @Override public void close() {}
}

public class RemoteSevenZReader {
    public static void listSevenZEntries(String sevenzUrl) throws Exception {
        URL url = new URL(sevenzUrl);
        try (SeekableByteChannel channel = new HttpSeekableByteChannel(url);
             SevenZFile sevenZFile = new SevenZFile(channel)) {
            SevenZArchiveEntry entry;
            while ((entry = sevenZFile.getNextEntry()) != null) {
                System.out.println(entry.getName());
            }
        }
    }
}

TAR格式

TAR无集中目录,需逐个请求512字节的文件头并判断结束标记:

import org.apache.commons.compress.archivers.tar.TarArchiveEntry;
import org.apache.commons.compress.archivers.tar.TarArchiveInputStream;
import java.net.HttpURLConnection;
import java.io.InputStream;
import java.net.URL;
import java.io.ByteArrayInputStream;

public class RemoteTarReader {
    public static void listTarEntries(String tarUrl) throws Exception {
        URL url = new URL(tarUrl);
        HttpURLConnection headConn = (HttpURLConnection) url.openConnection();
        headConn.setRequestMethod("HEAD");
        long fileSize = headConn.getContentLengthLong();
        headConn.disconnect();

        long offset = 0;
        while (offset < fileSize) {
            // 请求当前文件头
            HttpURLConnection conn = (HttpURLConnection) url.openConnection();
            conn.setRequestProperty("Range", "bytes=" + offset + "-" + (offset + 511));
            InputStream in = conn.getInputStream();
            byte[] headerBytes = in.readAllBytes();
            conn.disconnect();

            // 判断是否为全零结束标记
            boolean isEnd = true;
            for (byte b : headerBytes) {
                if (b != 0) {
                    isEnd = false;
                    break;
                }
            }
            if (isEnd) break;

            // 解析文件头并计算下一个偏移
            try (TarArchiveInputStream tis = new TarArchiveInputStream(new ByteArrayInputStream(headerBytes))) {
                TarArchiveEntry entry = tis.getNextTarEntry();
                if (entry != null) {
                    System.out.println(entry.getName());
                    offset += 512 + ((entry.getSize() + 511) / 512) * 512;
                } else {
                    offset += 512;
                }
            }
        }
    }
}

类似remotezip的通用工具封装

上面的HttpSeekableByteChannel是核心,它实现了SeekableByteChannel接口,可直接被Apache Commons Compress的SevenZFile、ZipFile等类使用,模拟本地文件的随机访问,无需下载完整文件。你可以基于此封装支持多格式的远程归档阅读器。

注意事项

  • 服务器兼容性:必须确保目标服务器支持HTTP范围请求,否则只能下载完整文件。
  • 格式限制:Apache Commons Compress对RAR支持有限,如需完整RAR支持需使用第三方商业库。
  • 边界处理:归档文件可能包含可变长度的注释(如ZIP的EOCD注释),解析时需考虑此类情况避免索引定位错误。
  • 性能优化:批量请求字节范围可减少HTTP连接数,提升效率。

内容的提问来源于stack exchange,提问作者Alper M.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 18:23:19