Java通过URLConnection/HttpsURLConnection访问Maven仓库返回403错误
问题背景
目标:当前仅作为POC,用于自动定期检索Maven仓库中的CVE标签。
通过浏览器、mvn命令均可正常访问Maven仓库,但使用Java代码访问时始终返回403。已尝试使用URLConnection、HttpsURLConnection发起请求,分别测试了携带/不携带GET请求方法、Content-Type、User-Agent、Accept请求头的场景,所有访问Maven仓库地址的请求均返回403状态码。同一段代码访问cve.mitre.org、nvd.nist.gov等其他站点可正常运行,但访问https://mvnrepository.com/artifact/log4j/apache-log4j-extras/1.2.17时请求失败。
URL为动态拼接生成,前缀固定为https://mvnrepository.com/artifact/,后续依次拼接构件的groupId、名称、版本号,最终生成合法访问地址,例如https://mvnrepository.com/artifact/log4j/apache-log4j-extras/1.2.17。
问题复现代码
System.setProperty("https.proxyHost", "xxxx"); System.setProperty("https.proxyPort", "xxxx"); String content = null; try { URL obj = new URL(address); HttpsURLConnection con = (HttpsURLConnection) obj.openConnection(); con.setRequestMethod("GET"); con.setRequestProperty("Content-Type", "application/json"); con.setRequestProperty("User-Agent", "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/65.0.3325.181 Safari/537.36"); con.setRequestProperty("Accept", "*/*"); con.connect(); BufferedReader br; if (con.getResponseCode() < 300) { br = new BufferedReader(new InputStreamReader(con.getInputStream(), StandardCharsets.UTF_8)); } else { br = new BufferedReader(new InputStreamReader(con.getErrorStream(), StandardCharsets.UTF_8)); } final StringBuilder sb = new StringBuilder(); String line; while ((line = br.readLine()) != null) { sb.append(line); } br.close();
解决方案
返回403是mvnrepository的反爬拦截导致的,代码存在两处明显的爬虫特征,直接触发了拦截规则:
- GET请求无请求体,你却携带了
Content-Type: application/json头,正常浏览器访问静态页面不会带这个头,属于脚本请求的典型特征。 - 你使用的Chrome 65版本User-Agent是2018年的老旧版本,已经被纳入公开爬虫UA特征库,会被直接判定为爬虫拦截。
按以下规则修改即可正常访问:
- 直接移除GET请求的
Content-Type请求头,无请求体时不需要传该字段。 - 替换为近期版本的浏览器UA,补全常规浏览器默认携带的
Accept-Language、Connection、Upgrade-Insecure-Requests头,模拟真实浏览器的请求特征。 - 控制请求频率,单IP请求间隔不要低于3秒,高频请求会被临时封禁IP,同样返回403。
修改后的核心请求代码参考:
System.setProperty("https.proxyHost", "xxxx"); System.setProperty("https.proxyPort", "xxxx"); String content = null; try { URL obj = new URL(address); HttpsURLConnection con = (HttpsURLConnection) obj.openConnection(); con.setRequestMethod("GET"); // 移除错误的Content-Type头 con.setRequestProperty("User-Agent", "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36"); con.setRequestProperty("Accept", "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8"); con.setRequestProperty("Accept-Language", "zh-CN,zh;q=0.9,en;q=0.8"); con.setRequestProperty("Connection", "keep-alive"); con.setRequestProperty("Upgrade-Insecure-Requests", "1"); con.connect(); BufferedReader br; // 测试阶段不要传Accept-Encoding头,避免额外处理gzip/br压缩流的解压逻辑 if (con.getResponseCode() < 300) { br = new BufferedReader(new InputStreamReader(con.getInputStream(), StandardCharsets.UTF_8)); } else { br = new BufferedReader(new InputStreamReader(con.getErrorStream(), StandardCharsets.UTF_8)); } final StringBuilder sb = new StringBuilder(); String line; while ((line = br.readLine()) != null) { sb.append(line); } br.close(); con.disconnect();
如果是长期做CVE检索的POC,建议直接对接Maven官方仓库的搜索接口获取元数据,不要爬取mvnrepository的前端页面,稳定性更高也不会触发反爬拦截。
内容的提问来源于stack exchange,提问作者res
相关产品推荐
相关产品推荐

