如何通过URL链接递归读取所有文件与文件夹并批量下载
解决方案:批量下载服务器目录下的所有文件
首先得明确一个核心前提:要批量下载目标目录/var/www/html/folders_and_files_that_i_need下的所有文件,服务器必须允许目录列表访问(比如Apache/Nginx开启了目录索引功能,访问该路径会返回包含所有文件/子目录的HTML页面)。如果服务器没有暴露目录结构,那你得先确认是否有其他方式获取文件列表(比如后端提供的API),否则无法自动遍历下载。
下面是基于你现有代码扩展的完整解决方案:
步骤1:扩展代码,添加目录遍历与批量下载逻辑
我们会新增三个核心功能,同时复用你原有的单文件下载逻辑:
- 解析服务器返回的目录HTML页面,提取所有文件和子目录链接
- 递归遍历子目录,保持本地保存路径与服务器目录结构一致
- 整合原有下载逻辑,实现批量下载
import org.jsoup.Jsoup; import org.jsoup.nodes.Document; import org.jsoup.nodes.Element; import org.jsoup.select.Elements; import java.io.*; import java.net.HttpURLConnection; import java.net.URL; import java.util.ArrayList; import java.util.List; public class BatchFileDownloader { private static final int BUFFER_SIZE = 4096; // 替换成你的服务器基础域名(比如http://your-server.com) private static final String BASE_SERVER_URL = "http://your-server-domain.com"; // 目标服务器目录路径 private static final String TARGET_SERVER_DIR = "/var/www/html/folders_and_files_that_i_need"; // 本地保存根目录(可自定义) private static final String LOCAL_SAVE_ROOT = "D:/downloads/target_files"; public static void main(String[] args) { try { // 启动递归下载 downloadDirectory(TARGET_SERVER_DIR, LOCAL_SAVE_ROOT); System.out.println("所有文件下载完成!"); } catch (IOException e) { e.printStackTrace(); } } /** * 递归下载服务器目录及其子目录下的所有文件 */ private static void downloadDirectory(String serverDirPath, String localSaveDir) throws IOException { // 创建本地保存目录(不存在则自动创建) File localDir = new File(localSaveDir); if (!localDir.exists()) { if (!localDir.mkdirs()) { throw new IOException("无法创建本地目录:" + localSaveDir); } } // 构造目录的访问URL String dirUrl = BASE_SERVER_URL + serverDirPath; // 获取目录下的所有文件和子目录链接 List<String> fileLinks = getDirectoryFileLinks(dirUrl); for (String link : fileLinks) { // 跳过上级目录链接(比如../) if (link.startsWith("../")) { continue; } // 构造完整的服务器文件/目录路径 String fullServerPath = serverDirPath + "/" + link; // 构造本地保存路径 String localPath = localSaveDir + File.separator + link; // 判断是文件还是子目录:链接以/结尾则为目录 if (link.endsWith("/")) { // 递归下载子目录 downloadDirectory(fullServerPath, localPath); } else { // 调用原有逻辑下载单个文件 downloadFile(BASE_SERVER_URL + fullServerPath, localSaveDir); } } } /** * 解析服务器目录HTML页面,提取所有文件和子目录链接 * 依赖Jsoup库解析HTML,需在项目中引入 */ private static List<String> getDirectoryFileLinks(String dirUrl) throws IOException { List<String> links = new ArrayList<>(); Document doc = Jsoup.connect(dirUrl).get(); // 选择目录页面中的所有有效链接(跳过当前目录./) Elements aElements = doc.select("a[href]"); for (Element a : aElements) { String href = a.attr("href"); if (!href.equals("./")) { links.add(href); } } return links; } // 你的原有单文件下载方法,保留核心逻辑并优化输出 public static void downloadFile(String fileURL, String saveDir) throws IOException { URL url = new URL(fileURL); HttpURLConnection httpConn = (HttpURLConnection) url.openConnection(); int responseCode = httpConn.getResponseCode(); if (responseCode == HttpURLConnection.HTTP_OK) { String fileName = ""; String disposition = httpConn.getHeaderField("Content-Disposition"); String contentType = httpConn.getContentType(); int contentLength = httpConn.getContentLength(); if (disposition != null) { int index = disposition.indexOf("filename="); if (index > 0) { fileName = disposition.substring(index + 10, disposition.length() - 1); } } else { fileName = fileURL.substring(fileURL.lastIndexOf("/") + 1); } System.out.println("正在下载: " + fileName); System.out.println("文件类型: " + contentType); System.out.println("文件大小: " + contentLength + " bytes"); InputStream inputStream = httpConn.getInputStream(); String saveFilePath = saveDir + File.separator + fileName; FileOutputStream outputStream = new FileOutputStream(saveFilePath); int bytesRead = -1; byte[] buffer = new byte[BUFFER_SIZE]; while ((bytesRead = inputStream.read(buffer)) != -1) { outputStream.write(buffer, 0, bytesRead); } outputStream.close(); inputStream.close(); System.out.println("✅ " + fileName + " 下载完成"); } else { System.out.println("❌ 无法下载文件,服务器返回HTTP状态码: " + responseCode); } httpConn.disconnect(); } }
步骤2:引入必要依赖
上面的代码使用Jsoup库解析HTML目录页面,如果你用Maven,需要在pom.xml中添加:
<dependency> <groupId>org.jsoup</groupId> <artifactId>jsoup</artifactId> <version>1.17.2</version> </dependency>
如果是Gradle,添加:
implementation 'org.jsoup:jsoup:1.17.2'
关键注意事项
- 服务器目录权限:如果访问目标目录返回403/404,说明服务器未开启目录索引,需要联系管理员开启,或者获取官方文件列表API。
- 认证处理:如果服务器需要登录,需在
HttpURLConnection中添加Cookie或Authorization头,比如:// 添加Cookie示例 httpConn.setRequestProperty("Cookie", "your-login-cookie"); // Basic认证示例 String auth = "username:password"; String encodedAuth = java.util.Base64.getEncoder().encodeToString(auth.getBytes()); httpConn.setRequestProperty("Authorization", "Basic " + encodedAuth); - 并发优化:文件数量较多时,可使用
ExecutorService线程池实现并行下载,提升效率。 - 跨平台兼容:代码中使用
File.separator保证Windows/Linux路径都能正常工作。
内容的提问来源于stack exchange,提问作者Julio Amorim
相关产品推荐
相关产品推荐

