You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Groovy/Grails无法完整解析HTML页面求助

解决HTML解析不完整、无法获取下载链接的问题

可能的原因

  • 目标页面依赖JavaScript动态渲染内容,直接通过XML解析器请求的是原始HTML,未包含JS加载的部分
  • 请求时未模拟浏览器请求头(如User-Agent、Cookie),服务器返回了不完整/精简版的HTML
  • 页面存在严重的HTML语法错误,导致Tagsoup/NekoHTML解析时提前终止

解决方案

1. 使用Jsoup解析(推荐,处理现代HTML更健壮)

Jsoup对有语法问题的HTML兼容性更好,支持CSS选择器直接定位下载链接,无需遍历所有元素。

步骤:

  • 添加Jsoup依赖到BuildConfig.groovy:
dependencies {
    compile 'org.jsoup:jsoup:1.17.2' // 适配Grails 2.5.6的Java 7+环境
}
  • 解析页面并提取下载链接的代码:
import org.jsoup.Jsoup
import org.jsoup.nodes.Document
import org.jsoup.nodes.Element

// 模拟浏览器请求头,确保拿到完整页面
def headers = [
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:109.0) Gecko/20100101 Firefox/115.0",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8"
]

// 获取页面文档
Document doc = Jsoup.connect("https://somewebsite")
                    .headers(headers)
                    .get()

// 用CSS选择器定位所有目标下载链接(匹配class和属性特征)
def downloadLinks = doc.select("a.p2n-btn-download[href][title='Download file']")

// 提取并保存链接
new File("myFile.txt").withWriter { writer ->
    downloadLinks.each { Element link ->
        String absoluteUrl = link.absUrl("href") // 自动转换相对链接为绝对链接
        writer.writeLine("下载链接:$absoluteUrl")
        writer.writeLine("元素原始HTML:${link.outerHtml()}")
        writer.writeLine("---")
    }
}

2. 处理动态渲染的页面(JS加载内容)

如果目标链接是通过JavaScript动态生成的,直接解析静态HTML无法获取,需要用无头浏览器模拟渲染:

使用HtmlUnit(无界面浏览器,适合后台处理)

  • 添加依赖:
dependencies {
    compile 'net.sourceforge.htmlunit:htmlunit:2.70.0' // 适配Java 7+环境
}
  • 代码示例:
import com.gargoylesoftware.htmlunit.WebClient
import com.gargoylesoftware.htmlunit.html.HtmlPage
import org.jsoup.Jsoup
import org.jsoup.nodes.Document

// 初始化WebClient,启用JS执行
WebClient webClient = new WebClient()
webClient.getOptions().setJavaScriptEnabled(true)
webClient.getOptions().setCssEnabled(false) // 关闭CSS渲染提升速度
webClient.getOptions().setThrowExceptionOnScriptError(false) // 忽略JS执行错误

try {
    // 获取渲染后的页面
    HtmlPage page = webClient.getPage("https://somewebsite")
    // 等待JS执行完成(根据页面加载速度调整等待时间,单位毫秒)
    webClient.waitForBackgroundJavaScript(5000)

    // 将渲染后的HTML转为Jsoup文档处理
    Document doc = Jsoup.parse(page.asXml())
    // 定位下载链接
    def downloadLinks = doc.select("a.p2n-btn-download[href][title='Download file']")

    // 保存结果
    new File("myFile.txt").withWriter { writer ->
        downloadLinks.each { link ->
            writer.writeLine("绝对链接:${link.absUrl('href')}")
            writer.writeLine("元素HTML:${link.outerHtml()}")
            writer.writeLine("---")
        }
    }
} finally {
    webClient.close()
}

3. 检查原始响应内容

如果以上方法仍无效,先确认请求到的HTML是否完整:

import java.net.URL
import java.net.HttpURLConnection

def url = new URL("https://somewebsite")
HttpURLConnection conn = (HttpURLConnection) url.openConnection()
conn.setRequestProperty("User-Agent", "Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:109.0) Gecko/20100101 Firefox/115.0")

new File("rawResponse.html").withOutputStream { out ->
    conn.inputStream.copyTo(out)
}

打开rawResponse.html查看是否包含目标下载链接,如果没有,说明服务器返回内容本身不完整,需要检查是否需要添加Cookie、登录认证或其他请求头。

内容的提问来源于stack exchange,提问作者Simon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 20:02:21