Java下载网络文件报403错误、下载为HTML或失败如何解决?
Java网络文件下载403异常/下载内容为HTML的解决办法
问题根因
目标服务器配置了反爬策略,默认
java.net.URL发起请求时携带的标识为Java/版本号的User-Agent会被服务器识别为非浏览器请求,因此要么返回403拒绝访问,要么返回跳转/验证页面的HTML内容,而非目标文件。
修复方案
- 改用
HttpURLConnection发起请求,手动配置请求头模拟浏览器行为 - 开启自动跟随重定向配置,解决部分文件下载需要跳转的问题
- 使用try-with-resources自动管理流资源,避免异常场景下的资源泄漏
修改后的完整代码
import java.io.File; import java.io.FileOutputStream; import java.io.IOException; import java.io.InputStream; import java.net.HttpURLConnection; import java.net.URL; public class URLReader { public static void copyURLToFile(URL url, File file) throws IOException { HttpURLConnection connection = (HttpURLConnection) url.openConnection(); // 配置请求头模拟Chrome浏览器请求,可替换为你自己浏览器的真实UA connection.setRequestProperty("User-Agent", "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"); connection.setRequestProperty("Referer", url.getProtocol() + "://" + url.getHost()); // 开启自动跟随重定向 connection.setInstanceFollowRedirects(true); connection.setConnectTimeout(15000); connection.setReadTimeout(60000); // 先校验响应状态码 int responseCode = connection.getResponseCode(); if (responseCode < 200 || responseCode >= 300) { throw new IOException("请求异常,响应码:" + responseCode); } // 使用try-with-resources自动关闭流,无需手动关闭 try (InputStream input = connection.getInputStream(); FileOutputStream output = new FileOutputStream(file)) { if (file.exists()) { if (file.isDirectory()) throw new IOException("File '" + file + "' is a directory"); if (!file.canWrite()) throw new IOException("File '" + file + "' cannot be written"); } else { File parent = file.getParentFile(); if ((parent != null) && (!parent.exists()) && (!parent.mkdirs())) { throw new IOException("File '" + file + "' could not be created"); } } byte[] buffer = new byte[4096]; int n = 0; while (-1 != (n = input.read(buffer))) { output.write(buffer, 0, n); } System.out.println("File '" + file + "' downloaded successfully!"); } finally { connection.disconnect(); } } public static void main(String[] args) { //URL pointing to the file String sUrl = "https://cdn-103.anonfiles.com/11HfTeD0uc/69e1a574-1630531823/main.exe"; //File where to be downloaded File file = new File("C:\\Users\\mader\\Desktop\\main.exe"); try { URL url = new URL(sUrl); URLReader.copyURLToFile(url, file); } catch (IOException ioEx) { ioEx.printStackTrace(); } } }
额外注意事项
- 如果修改后依然返回403,可打开浏览器访问目标下载地址,按F12打开开发者工具,在「网络」标签下找到对应下载请求的
User-Agent和Cookie字段,复制替换到代码的请求头配置中即可 - 部分站点有下载频次限制,短时间多次请求也可能触发403,可间隔一段时间后重试
内容的提问来源于stack exchange,提问作者SSLWasTaken
相关产品推荐
相关产品推荐

