You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spring应用如何获取指定URL页面实际显示内容而非HTML源码?

问题本质

普通的URL内容拉取方法本质是发送裸HTTP请求,仅能拿到服务端返回的静态HTML源码,不会执行页面内的JavaScript代码,所有靠JS动态插入到DOM的可见内容都无法直接获取。要拿到最终渲染完成的页面展示值,必须使用支持JS执行、DOM渲染的无头浏览器组件,模拟真实浏览器的完整加载流程。

适配Spring项目的落地方案

方案1:HtmlUnit(轻量无额外依赖)

纯Java实现的无头浏览器,不需要安装额外的浏览器程序,直接引入依赖即可使用,适合逻辑简单的页面抓取场景。

  1. 引入Maven依赖
<dependency>
    <groupId>net.sourceforge.htmlunit</groupId>
    <artifactId>htmlunit</artifactId>
    <version>2.70.0</version>
</dependency>
  1. 代码实现
import com.gargoylesoftware.htmlunit.BrowserVersion;
import com.gargoylesoftware.htmlunit.WebClient;
import com.gargoylesoftware.htmlunit.html.HtmlPage;

public class PageContentFetcher {
    public String getRenderedPageContent(String targetUrl) throws Exception {
        try (WebClient webClient = new WebClient(BrowserVersion.CHROME)) {
            // 关闭非必要能力提升抓取效率
            webClient.getOptions().setCssEnabled(false);
            webClient.getOptions().setThrowExceptionOnScriptError(false);
            webClient.getOptions().setThrowExceptionOnFailingStatusCode(false);
            // 设置等待后台JS执行的超时时间,单位毫秒
            webClient.waitForBackgroundJavaScript(10000);
            HtmlPage page = webClient.getPage(targetUrl);
            // 直接提取页面所有可见文本,自动过滤标签、脚本内容
            return page.asNormalizedText();

            // 如果需要提取指定元素内容,可使用选择器定位
            // String specificContent = page.querySelector("#target-element-id").asNormalizedText();
        }
    }
}
  • 优劣势:集成成本低、内存占用小;对使用最新ES语法、复杂前端框架的页面兼容性稍差,可能出现JS执行失败的情况。

方案2:Selenium + 无头Chrome(兼容性拉满)

驱动真实Chrome/Edge浏览器以无头模式运行,渲染效果和用户手动打开浏览器完全一致,适合逻辑复杂、前端渲染重的页面。

  1. 引入Maven依赖
<dependency>
    <groupId>org.seleniumhq.selenium</groupId>
    <artifactId>selenium-java</artifactId>
    <version>4.15.0</version>
</dependency>

注意:运行环境需要安装对应版本的Chrome/Edge浏览器,Selenium 4+版本内置驱动管理器,无需手动下载浏览器驱动。

  1. 代码实现
import org.openqa.selenium.By;
import org.openqa.selenium.chrome.ChromeDriver;
import org.openqa.selenium.chrome.ChromeOptions;
import java.time.Duration;

public class PageContentFetcher {
    public String getRenderedPageContent(String targetUrl) {
        ChromeOptions options = new ChromeOptions();
        // 开启无头模式,不弹出可视化窗口
        options.addArguments("--headless=new");
        // 适配服务器运行环境的配置项
        options.addArguments("--no-sandbox", "--disable-gpu", "--disable-dev-shm-usage");
        try (ChromeDriver driver = new ChromeDriver(options)) {
            driver.manage().timeouts().pageLoadTimeout(Duration.ofSeconds(15));
            driver.get(targetUrl);
            // 异步加载页面可配置显式等待,等目标元素加载完成再取值
            // WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(10));
            // wait.until(ExpectedConditions.visibilityOfElementLocated(By.id("target-element-id")));
            // 提取body下所有可见文本
            return driver.findElement(By.tagName("body")).getText();
        }
    }
}
  • 优劣势:渲染结果和真实浏览器完全一致,兼容性极强;内存占用更高,服务器部署需要额外适配浏览器运行环境。
避坑提示
  • 不要使用getPageSource()、asXml()这类方法获取内容,这类方法返回的是原始DOM结构,和你之前拿到的静态HTML没有本质区别,提取可见文本必须用asNormalizedText()、getText()这类专门的文本提取方法。
  • 针对存在延迟加载、滚动加载的页面,不要页面刚加载完成就提取内容,加显式等待逻辑确认目标内容渲染完成后再取值,避免拿到不完整的结果。
  • 遇到有反爬校验的站点,可在请求头中补充真实的User-Agent、Cookie等信息,模拟普通用户的请求特征。

内容的提问来源于stack exchange,提问作者Chilled Shark

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 18:51:23