You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Java无头浏览器HtmlUnit获取Crystal Server跳转后的报表PDF

HtmlUnit获取Crystal Server PDF报表的重定向问题解决办法

刚接触HtmlUnit,需求是调用Crystal Server获取报表,目前用该服务器暴露的Restful API,但Crystal Server没有直接获取报表的API。从某API端点拿到最终链接后,常规浏览器打开会经过三次左右重定向才加载PDF文档。试了用HtmlUnit模拟浏览器行为,代码如下:

try (final WebClient webClient = new WebClient()) {
    webClient.getOptions().setJavaScriptEnabled(true);
    webClient.getOptions().setThrowExceptionOnScriptError(false);
    webClient.getOptions().setThrowExceptionOnFailingStatusCode(false);
    webClient.getOptions().setRedirectEnabled(true);
    htmlPage = webClient.getPage(linkString);
}

但现在只能完成第二次重定向,拿不到最终的PDF文档。想知道怎么获取最终的文档页面,是不是得捕获结果后用新WebClient发第三次请求,或者有更简单的实现方式?

几个可行的解决办法

1. 直接获取原始响应,处理二进制PDF流

HtmlUnit默认会把响应解析成HTML页面,但最终的PDF是二进制流,这可能是它没法自动处理最后一次重定向的原因。可以跳过HtmlPage解析,直接拿WebResponse来处理:

try (final WebClient webClient = new WebClient()) {
    webClient.getOptions().setJavaScriptEnabled(true);
    webClient.getOptions().setThrowExceptionOnScriptError(false);
    webClient.getOptions().setThrowExceptionOnFailingStatusCode(false);
    webClient.getOptions().setRedirectEnabled(true);
    
    // 发起请求,直接获取WebResponse
    WebResponse webResponse = webClient.getPage(linkString).getWebResponse();
    
    // 判断是不是PDF类型的响应
    if ("application/pdf".equals(webResponse.getContentType())) {
        // 读取二进制数据,保存成文件
        byte[] pdfBytes = webResponse.getContentAsBytes();
        Files.write(Paths.get("report.pdf"), pdfBytes);
    } else {
        // 如果还没到PDF,手动拿Location头继续请求
        String redirectUrl = webResponse.getResponseHeaderValue("Location");
        if (redirectUrl != null) {
            webResponse = webClient.getPage(redirectUrl).getWebResponse();
            if ("application/pdf".equals(webResponse.getContentType())) {
                byte[] pdfBytes = webResponse.getContentAsBytes();
                Files.write(Paths.get("report.pdf"), pdfBytes);
            }
        }
    }
}

2. 自定义WebConnection拦截PDF响应

HtmlUnit对非HTML内容的处理有限,可以通过自定义WebConnection来拦截PDF响应,直接处理:

try (final WebClient webClient = new WebClient()) {
    webClient.getOptions().setJavaScriptEnabled(true);
    webClient.getOptions().setThrowExceptionOnScriptError(false);
    webClient.getOptions().setThrowExceptionOnFailingStatusCode(false);
    webClient.getOptions().setRedirectEnabled(true);
    
    // 包装WebConnection,拦截PDF类型的响应
    webClient.setWebConnection(new WebConnectionWrapper(webClient) {
        @Override
        public WebResponse getResponse(WebRequest request) throws IOException {
            WebResponse response = super.getResponse(request);
            // 是PDF就直接返回,不解析成HTML
            if ("application/pdf".equals(response.getContentType())) {
                return response;
            }
            return response;
        }
    });
    
    WebResponse finalResponse = webClient.getPage(linkString).getWebResponse();
    if ("application/pdf".equals(finalResponse.getContentType())) {
        byte[] pdfBytes = finalResponse.getContentAsBytes();
        Files.write(Paths.get("report.pdf"), pdfBytes);
    }
}

3. 关闭自动重定向,手动跟踪每一步

如果自动重定向在第三次失效,干脆关掉自动重定向,自己处理每一次的Location头:

try (final WebClient webClient = new WebClient()) {
    webClient.getOptions().setJavaScriptEnabled(true);
    webClient.getOptions().setThrowExceptionOnScriptError(false);
    webClient.getOptions().setThrowExceptionOnFailingStatusCode(false);
    // 关闭自动重定向
    webClient.getOptions().setRedirectEnabled(false);
    
    String currentUrl = linkString;
    WebResponse response;
    
    // 手动跟踪最多5次重定向,防止无限循环
    for (int i = 0; i < 5; i++) {
        response = webClient.getPage(currentUrl).getWebResponse();
        int statusCode = response.getStatusCode();
        
        // 检查是不是3xx重定向状态码
        if (statusCode >= 300 && statusCode < 400) {
            currentUrl = response.getResponseHeaderValue("Location");
            // 处理相对路径的URL
            if (!currentUrl.startsWith("http")) {
                URL baseUrl = new URL(currentUrl);
                currentUrl = new URL(baseUrl, currentUrl).toString();
            }
        } else {
            // 不是重定向,检查是不是PDF
            if ("application/pdf".equals(response.getContentType())) {
                byte[] pdfBytes = response.getContentAsBytes();
                Files.write(Paths.get("report.pdf"), pdfBytes);
            }
            break;
        }
    }
}

4. 检查会话Cookie是否正常保留

Crystal Server可能依赖会话Cookie来验证请求,HtmlUnit默认会自动管理Cookie,但可以手动确认一下:

// 打印当前所有Cookie,看会话相关的Cookie是否存在
Set<Cookie> cookies = webClient.getCookieManager().getCookies();
for (Cookie cookie : cookies) {
    System.out.println(cookie.getName() + ": " + cookie.getValue());
}

如果发现关键Cookie丢失,就从之前调用Restful API时的响应里手动提取并添加到WebClient的Cookie管理器中。

内容的提问来源于stack exchange,提问作者Munny

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 14:34:54