HtmlUnit解析ipinfo.io返回空列表的问题原因及解决方法
问题描述
我尝试从ipinfo.io网站获取IP相关数据字符串,使用HtmlUnit解析页面时,返回的列表为空,但该网站对应元素的实际内容并非空值。
我的Java代码如下:
LogFactory.getFactory().setAttribute("org.apache.commons.logging.Log", "org.apache.commons.logging.impl.NoOpLog"); java.util.logging.Logger.getLogger("com.gargoylesoftware").setLevel(Level.OFF); java.util.logging.Logger.getLogger("org.apache.commons.httpclient").setLevel(Level.OFF); final WebClient webClient = new WebClient(BrowserVersion.EDGE); webClient.getOptions().setThrowExceptionOnScriptError(false); webClient.getOptions().setAppletEnabled(true); webClient.getOptions().setCssEnabled(true); webClient.getOptions().setJavaScriptEnabled(false); webClient.getOptions().setFetchPolyfillEnabled(true); webClient.getCookieManager().setCookiesEnabled(true); webClient.getOptions().setUseInsecureSSL(true); final HtmlPage page1 = webClient.getPage("https://ipinfo.io/"); HtmlForm form = page1.getForms().get(0); List<DomElement> elements = StreamSupport.stream(form.getChildElements().spliterator(), false).collect(Collectors.toList()); HtmlTextInput textField = form.getInputByValue(""); textField.setValueAttribute("8.8.4.4"); HtmlPage page2 = elements.get(1).click(); HtmlDivision div = (HtmlDivision) page2.getElementById("tryit-data"); HtmlUnorderedList list = (HtmlUnorderedList) StreamSupport.stream(div.getChildElements().spliterator(), false).collect(Collectors.toList()).get(0); System.out.println(list.asXml());
运行代码后,输出的列表内容为空;但网站实际显示8.8.4.4的完整IP信息列表。
问题原因
- JavaScript禁用导致动态内容无法加载:ipinfo.io的查询结果是通过JavaScript动态渲染的,代码中设置了
webClient.getOptions().setJavaScriptEnabled(false),点击查询按钮后,页面无法执行JS获取并渲染IP数据,因此返回的元素内容为空。 - 未等待页面异步加载完成:即使启用JS,点击操作后页面需要时间加载异步数据,直接获取元素会拿到未更新的DOM结构。
- 元素定位不稳定:通过
elements.get(1)定位点击按钮,依赖表单子元素的顺序,一旦页面结构微调就会失效。
修复方案
1. 启用JavaScript并设置等待逻辑
针对页面动态渲染的特性,启用JS并等待异步请求完成:
LogFactory.getFactory().setAttribute("org.apache.commons.logging.Log", "org.apache.commons.logging.impl.NoOpLog"); java.util.logging.Logger.getLogger("com.gargoylesoftware").setLevel(Level.OFF); java.util.logging.Logger.getLogger("org.apache.commons.httpclient").setLevel(Level.OFF); final WebClient webClient = new WebClient(BrowserVersion.EDGE); webClient.getOptions().setThrowExceptionOnScriptError(false); webClient.getOptions().setAppletEnabled(true); webClient.getOptions().setCssEnabled(true); // 启用JavaScript webClient.getOptions().setJavaScriptEnabled(true); webClient.getOptions().setFetchPolyfillEnabled(true); webClient.getCookieManager().setCookiesEnabled(true); webClient.getOptions().setUseInsecureSSL(true); // 设置等待JS执行完成的时间 webClient.waitForBackgroundJavaScript(5000); final HtmlPage page1 = webClient.getPage("https://ipinfo.io/"); HtmlForm form = page1.getForms().get(0); // 通过name定位输入框,比按value定位更稳定 HtmlTextInput textField = form.getInputByName("ip"); textField.setValueAttribute("8.8.4.4"); // 通过按钮文本定位查询按钮,避免依赖元素顺序 HtmlButton submitBtn = form.getButtonByValue("Lookup"); HtmlPage page2 = submitBtn.click(); // 等待页面异步请求完成 webClient.waitForBackgroundJavaScript(5000); HtmlDivision div = (HtmlDivision) page2.getElementById("tryit-data"); // 用XPath直接定位列表元素 HtmlUnorderedList list = (HtmlUnorderedList) div.getFirstByXPath("./ul"); if (list != null) { System.out.println(list.asXml()); } else { System.out.println("未找到目标列表元素"); }
2. 直接调用API替代页面解析(更高效可靠)
ipinfo.io提供公开API,直接请求API无需处理JS渲染,稳定性和效率更高:
import java.io.BufferedReader; import java.io.InputStreamReader; import java.net.HttpURLConnection; import java.net.URL; public class IpInfoApi { public static void main(String[] args) throws Exception { String ip = "8.8.4.4"; URL url = new URL("https://ipinfo.io/" + ip + "/json"); HttpURLConnection conn = (HttpURLConnection) url.openConnection(); conn.setRequestMethod("GET"); BufferedReader in = new BufferedReader(new InputStreamReader(conn.getInputStream())); String inputLine; StringBuilder response = new StringBuilder(); while ((inputLine = in.readLine()) != null) { response.append(inputLine); } in.close(); System.out.println(response.toString()); } }
关键修复点说明
- 启用
JavaScriptEnabled(true),确保动态内容能正常渲染。 - 使用
waitForBackgroundJavaScript()等待异步请求完成,避免获取未更新的DOM。 - 替换不稳定的元素定位方式,通过
name、XPath或按钮文本定位元素,降低页面结构变化带来的影响。 - 优先选择API调用,页面解析易受网站结构调整影响,API则更稳定且响应更快。
内容的提问来源于stack exchange,提问作者komla3
相关产品推荐
相关产品推荐

