Java中短链接展开异常:部分链接返回null的问题求助
解决短链接展开时无Location头的问题
你的问题很常见——有些短链接服务并不是通过HTTP 3xx重定向(也就是Location响应头)来跳转,而是返回一个包含自动跳转逻辑的HTML页面(比如通过<meta>刷新标签或JavaScript代码)。当状态码是200时,你的代码自然拿不到Location头,所以返回null。
解决方案思路
我们需要分两种情况处理:
- 当响应是3xx重定向时,直接提取
Location头(保持你原有的逻辑); - 当响应是200时,解析返回的HTML内容,查找页面中的自动跳转指令,提取目标URL。
最方便的HTML解析工具是Jsoup,我们可以用它来快速定位跳转标签。
修改后的代码示例
首先,你需要引入Jsoup依赖(如果用Maven,在pom.xml中添加):
<dependency> <groupId>org.jsoup</groupId> <artifactId>jsoup</artifactId> <version>1.17.2</version> <!-- 用最新稳定版即可 --> </dependency>
然后修改你的ExpandUrl类:
import java.io.IOException; import java.net.HttpURLConnection; import java.net.Proxy; import java.net.URL; import org.jsoup.Jsoup; import org.jsoup.nodes.Document; import org.jsoup.nodes.Element; public class ExpandUrl { public static void main(String[] args) throws IOException { String shortenedUrl = "9mglq8"; // 测试有问题的短链接 String expandedURL = ExpandUrl.expand(shortenedUrl); System.out.println("expanded url is : " + expandedURL); } public static String expand(String shortenedUrl) throws IOException { // 确保短链接是完整的URL(如果用户只给了哈希,需要补全对应服务的前缀) if (!shortenedUrl.startsWith("http")) { // 这里根据实际短链接服务补全前缀,比如bit.ly的话就是https://bit.ly/ shortenedUrl = "https://bit.ly/" + shortenedUrl; } URL url = new URL(shortenedUrl); HttpURLConnection httpURLConnection = (HttpURLConnection) url.openConnection(Proxy.NO_PROXY); httpURLConnection.setInstanceFollowRedirects(false); // 禁止自动跳转 int status = httpURLConnection.getResponseCode(); System.out.println("status is " + status); // 情况1:3xx重定向,直接取Location头 if (status >= 300 && status < 400) { String expandedURL = httpURLConnection.getHeaderField("Location"); httpURLConnection.disconnect(); // 处理相对URL if (expandedURL != null && !expandedURL.startsWith("http")) { expandedURL = new URL(url, expandedURL).toString(); } return expandedURL; } // 情况2:200响应,解析HTML找跳转 else if (status == 200) { Document doc = Jsoup.parse(httpURLConnection.getInputStream(), "UTF-8", url.toString()); // 查找meta refresh标签(最常见的自动跳转方式) Element metaRefresh = doc.selectFirst("meta[http-equiv=refresh]"); if (metaRefresh != null) { String content = metaRefresh.attr("content"); // content格式通常是 "0; URL=https://example.com" if (content.contains("URL=")) { String targetUrl = content.split("URL=")[1].trim(); // 处理相对URL转绝对URL if (!targetUrl.startsWith("http")) { targetUrl = new URL(url, targetUrl).toString(); } httpURLConnection.disconnect(); return targetUrl; } } // 如果没有meta标签,查找JavaScript跳转(比如window.location.href) for (Element script : doc.select("script")) { String scriptContent = script.html(); if (scriptContent.contains("window.location") || scriptContent.contains("location.href")) { // 基础提取URL,复杂场景可替换为正则匹配 int startIndex = scriptContent.indexOf("\""); int endIndex = scriptContent.indexOf("\"", startIndex + 1); if (startIndex != -1 && endIndex != -1) { String targetUrl = scriptContent.substring(startIndex + 1, endIndex); if (!targetUrl.startsWith("http")) { targetUrl = new URL(url, targetUrl).toString(); } httpURLConnection.disconnect(); return targetUrl; } } } // 如果都找不到跳转指令,返回原短链接或null(根据需求调整) httpURLConnection.disconnect(); return shortenedUrl; } // 其他状态码,返回原链接或null else { httpURLConnection.disconnect(); return shortenedUrl; } } }
关键说明
- 补全短链接前缀:你之前用的
"2q3GMg0"这类哈希值,实际请求时需要补全对应短链接服务的完整域名前缀(比如https://bit.ly/),否则URL类会抛出异常,代码里已经添加了前缀补全逻辑,你可以根据实际使用的短链接服务调整。 - Meta标签解析:大部分这类短链接会用
<meta http-equiv="refresh" content="0; URL=目标地址">实现自动跳转,Jsoup可以快速定位并提取目标URL。 - JavaScript跳转处理:部分服务会用JS代码跳转,代码里做了基础的URL提取,复杂场景可以用更精准的正则表达式匹配。
- 相对URL处理:提取到的URL可能是相对路径,需要用
new URL(url, targetUrl)转换为绝对URL。
内容的提问来源于stack exchange,提问作者MONU KUMAR
相关产品推荐
相关产品推荐

