如何在Selenium 3.141.59中实现HTML转PDF?求替代方案
可行替代方案
以下是针对Selenium 3环境下实现页面转PDF的几种稳定方案:
方案1:通过ChromeDriver内部API调用DevTools命令
Selenium 3的ChromeDriver虽未暴露executeCdpCommand方法,但可以通过反射调用其内部的DevTools连接来发送Page.printToPDF命令:
import org.openqa.selenium.chrome.ChromeDriver; import com.google.common.collect.ImmutableMap; import java.lang.reflect.Method; import java.util.Map; public class PdfPrinter { public static byte[] printToPdf(ChromeDriver driver) throws Exception { // 获取ChromeDriver内部的命令发送方法 Method sendCommandMethod = ChromeDriver.class.getDeclaredMethod( "sendChromeCommand", String.class, Map.class ); sendCommandMethod.setAccessible(true); // 构造打印参数(可按需添加横向、背景打印等配置) Map<String, Object> params = ImmutableMap.of( "landscape", false, "printBackground", true ); // 发送Page.printToPDF命令 Map<String, Object> result = (Map<String, Object>) sendCommandMethod.invoke( driver, "Page.printToPDF", params ); // 解析返回的base64编码PDF数据并转为字节数组 String pdfBase64 = (String) result.get("data"); return java.util.Base64.getDecoder().decode(pdfBase64); } }
使用时直接调用printToPdf(driver)获取PDF字节数组,后续可写入文件保存。
方案2:提取页面HTML后用第三方库生成PDF
若DevTools调用方式受限,可先获取页面完整HTML(含渲染后样式),再用Flying Saucer或iText等工具将HTML转为PDF:
步骤1:获取页面完整HTML
// 获取基础页面源码,或执行JS获取包含动态内容的完整HTML String fullHtml = (String) driver.executeScript("return document.documentElement.outerHTML;");
步骤2:用Flying Saucer生成PDF
添加Maven依赖(若使用Maven):
<dependency> <groupId>org.xhtmlrenderer</groupId> <artifactId>flying-saucer-pdf</artifactId> <version>9.1.22</version> </dependency>
生成PDF的代码:
import org.xhtmlrenderer.pdf.ITextRenderer; import java.io.FileOutputStream; import java.io.StringReader; public class HtmlToPdf { public static void convert(String html, String outputPath) throws Exception { ITextRenderer renderer = new ITextRenderer(); renderer.setDocumentFromString(new StringReader(html)); renderer.layout(); try (FileOutputStream fos = new FileOutputStream(outputPath)) { renderer.createPDF(fos); } } }
注意:该方式需确保页面CSS样式能被Flying Saucer正确解析,复杂现代页面可能需要额外处理样式引用。
方案3:利用Chrome命令行工具打印PDF
通过命令行启动Chrome,加载Selenium会话保存的Cookie,直接打印目标页面为PDF:
步骤1:保存Selenium中的Cookie到文件
import org.openqa.selenium.Cookie; import java.io.FileWriter; import java.io.IOException; import java.util.Set; public class CookieSaver { public static void saveCookies(ChromeDriver driver, String cookiePath) throws IOException { Set<Cookie> cookies = driver.manage().getCookies(); try (FileWriter writer = new FileWriter(cookiePath)) { writer.write("["); boolean first = true; for (Cookie cookie : cookies) { if (!first) writer.write(","); writer.write(String.format( "{\"name\":\"%s\",\"value\":\"%s\",\"domain\":\"%s\",\"path\":\"%s\"}", cookie.getName(), cookie.getValue(), cookie.getDomain(), cookie.getPath() )); first = false; } writer.write("]"); } } }
步骤2:执行Chrome命令行打印PDF
import java.io.IOException; public class ChromePdfPrinter { public static void print(String url, String cookiePath, String outputPath) throws IOException, InterruptedException { String chromePath = "path/to/chrome.exe"; // 根据操作系统调整Chrome路径 ProcessBuilder pb = new ProcessBuilder( chromePath, "--headless", "--disable-gpu", "--no-sandbox", "--load-storage=" + cookiePath, "--print-to-pdf=" + outputPath, url ); pb.inheritIO().start().waitFor(); } }
该方式适合无需实时交互的场景,需确保Chrome版本与Selenium使用的ChromeDriver版本兼容。
内容的提问来源于stack exchange,提问作者sysOut
相关产品推荐
相关产品推荐

