如何解决Selenium+Java无法截取Google Drive PDF全页的问题
解决Selenium截取Google Drive PDF多页截图的翻页失效问题
问题描述
使用Java+Selenium实现Google Drive PDF多页截图时,仅能截取第一页,后续翻页操作无效。测试链接:https://drive.google.com/file/d/1optp32jv20rvyiSBcCdI_a7C1jRR1UT4/view
原代码
package googleDriveScreenshotTest; import org.openqa.selenium.*; import org.openqa.selenium.chrome.ChromeDriver; import org.openqa.selenium.io.FileHandler; import org.openqa.selenium.support.ui.ExpectedConditions; import org.openqa.selenium.support.ui.WebDriverWait; import java.io.File; import java.io.IOException; import java.time.Duration; import java.util.Scanner; public class PdfScreenShotTest2{ public static void main(String[] args) { Scanner scanner = new Scanner(System.in); // Get user inputs with validation String pdfUrl = getUserInput(scanner, "Enter Google Drive PDF link: "); String folderName = getUserInput(scanner, "Enter folder name to save images: "); int totalPages = getValidPageCount(scanner); // Setup WebDriver WebDriver driver = new ChromeDriver(); try { // Create directory with error handling File screenshotDir = createScreenshotDirectory(folderName); // Navigate to PDF driver.get(pdfUrl); driver.manage().window().maximize(); // Wait for document to load WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(15)); WebElement body = wait.until(ExpectedConditions.elementToBeClickable(By.tagName("body"))); // Click anywhere on page body.click(); // Take screenshots of each page for (int pageNumber = 1; pageNumber <= totalPages; pageNumber++) { File screenshot = ((TakesScreenshot) driver).getScreenshotAs(OutputType.FILE); File destination = new File(screenshotDir + "/page_" + pageNumber + ".png"); try { FileHandler.copy(screenshot, destination); System.out.println("Saved: " + destination.getAbsolutePath()); if (pageNumber < totalPages) { body.sendKeys(Keys.PAGE_DOWN); Thread.sleep(2000); // Allow time for scrolling } } catch (IOException e) { System.err.println("Error saving screenshot for page " + pageNumber + ": " + e.getMessage()); } } } catch (Exception e) { System.err.println("An error occurred: " + e.getMessage()); } finally { driver.quit(); scanner.close(); } } private static String getUserInput(Scanner scanner, String prompt) { while (true) { System.out.print(prompt); String input = scanner.nextLine().trim(); if (!input.isEmpty()) { return input; } System.out.println("Input cannot be empty. Please try again."); } } private static int getValidPageCount(Scanner scanner) { while (true) { try { System.out.print("Enter number of pages in PDF: "); int count = scanner.nextInt(); scanner.nextLine(); // Consume newline if (count > 0) { return count; } System.out.println("Please enter a positive number."); } catch (Exception e) { System.err.println("Invalid input. Please enter a valid number."); scanner.nextLine(); // Clear invalid input } } } private static File createScreenshotDirectory(String folderName) throws IOException { String downloadPath = System.getProperty("user.home") + "/Downloads/" + folderName; File folder = new File(downloadPath); if (!folder.exists()) { if (!folder.mkdirs()) { throw new IOException("Failed to create directory: " + downloadPath); } } return folder; } }
问题分析
- 焦点未定位到PDF阅读器:Google Drive的PDF阅读器是独立内嵌组件,直接对
body元素发送按键事件无法传递到阅读器内部,导致翻页指令无效。 - 等待机制不可靠:
Thread.sleep属于固定等待,无法确保页面完全切换,可能在页面未加载完成时就截图,导致重复截取第一页。
解决方案
修改代码,聚焦到PDF阅读器的内容区域,使用显式等待确认翻页完成,并替换固定等待为动态等待:
package googleDriveScreenshotTest; import org.openqa.selenium.*; import org.openqa.selenium.chrome.ChromeDriver; import org.openqa.selenium.io.FileHandler; import org.openqa.selenium.support.ui.ExpectedConditions; import org.openqa.selenium.support.ui.WebDriverWait; import java.io.File; import java.io.IOException; import java.time.Duration; import java.util.Scanner; public class PdfScreenShotTest2{ public static void main(String[] args) { Scanner scanner = new Scanner(System.in); // Get user inputs with validation String pdfUrl = getUserInput(scanner, "Enter Google Drive PDF link: "); String folderName = getUserInput(scanner, "Enter folder name to save images: "); int totalPages = getValidPageCount(scanner); // Setup WebDriver WebDriver driver = new ChromeDriver(); try { // Create directory with error handling File screenshotDir = createScreenshotDirectory(folderName); // Navigate to PDF driver.get(pdfUrl); driver.manage().window().maximize(); WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(20)); // 等待PDF阅读器加载完成,定位到内容容器 WebElement pdfContainer = wait.until(ExpectedConditions.visibilityOfElementLocated(By.cssSelector(".pdf-viewer-container"))); // 聚焦到PDF内容区域 pdfContainer.click(); // 记录初始页面的特征,用于确认翻页 String previousPageText = pdfContainer.getText(); // Take screenshots of each page for (int pageNumber = 1; pageNumber <= totalPages; pageNumber++) { File screenshot = ((TakesScreenshot) driver).getScreenshotAs(OutputType.FILE); File destination = new File(screenshotDir + "/page_" + pageNumber + ".png"); try { FileHandler.copy(screenshot, destination); System.out.println("Saved: " + destination.getAbsolutePath()); if (pageNumber < totalPages) { // 发送翻页指令 pdfContainer.sendKeys(Keys.PAGE_DOWN); // 等待页面切换,直到内容变化 wait.until(driver -> !pdfContainer.getText().equals(previousPageText)); // 更新当前页面文本特征 previousPageText = pdfContainer.getText(); } } catch (IOException e) { System.err.println("Error saving screenshot for page " + pageNumber + ": " + e.getMessage()); } } } catch (Exception e) { System.err.println("An error occurred: " + e.getMessage()); e.printStackTrace(); } finally { driver.quit(); scanner.close(); } } private static String getUserInput(Scanner scanner, String prompt) { while (true) { System.out.print(prompt); String input = scanner.nextLine().trim(); if (!input.isEmpty()) { return input; } System.out.println("Input cannot be empty. Please try again."); } } private static int getValidPageCount(Scanner scanner) { while (true) { try { System.out.print("Enter number of pages in PDF: "); int count = scanner.nextInt(); scanner.nextLine(); // Consume newline if (count > 0) { return count; } System.out.println("Please enter a positive number."); } catch (Exception e) { System.err.println("Invalid input. Please enter a valid number."); scanner.nextLine(); // Clear invalid input } } } private static File createScreenshotDirectory(String folderName) throws IOException { String downloadPath = System.getProperty("user.home") + "/Downloads/" + folderName; File folder = new File(downloadPath); if (!folder.exists()) { if (!folder.mkdirs()) { throw new IOException("Failed to create directory: " + downloadPath); } } return folder; } }
关键修改点
- 定位PDF阅读器容器:通过
By.cssSelector(".pdf-viewer-container")定位到Google Drive的PDF内容区域,确保焦点在阅读器内。 - 动态等待翻页完成:通过对比页面文本内容的变化,确认翻页完成,替换了不可靠的
Thread.sleep。 - 聚焦后发送按键:先点击PDF容器获取焦点,再发送
PAGE_DOWN指令,确保按键事件传递到阅读器。
内容的提问来源于stack exchange,提问作者Momo Maurice
相关产品推荐
相关产品推荐

