Java Selenium问题:无法从Chrome下载图片并保存为.png文件
问题描述
我想通过Chrome浏览器下载图片并保存为.png文件,但运行代码后,Download目录里生成的文件仅417字节,用Paint或其他程序打开时会报错:
"This is not a valid bitmap file or its format is not currently supported."
排查后发现,img元素的src属性返回的是完整HTTP页面的URL,而非直接指向图片的URL。我添加了以下代码查看内容类型:
String contentType = connection.getContentType(); System.out.println("Content Type: " + contentType);
返回的contentType为"text/html"; charset=ISO-8859-1,说明下载的是整个HTML页面,而非目标图片。
核心代码如下:
WebElement image = webDriver.findElement(By.xpath("//img")); String imageUrl = image.getAttribute("src"); String directory = System.getProperty("user.dir") + File.separator + "Downloads" + File.separator; try { URL url = new URL(imageURL); URLConnection connection = url.openConnection(); InputStream inputStream = connection.getInputStream(); OutputStream outputStream = new FileOutputStream(directory+"downloadedImage.png"); byte[] buffer = new byte[2048]; int length; while ((length = inputStream.read(buffer)) != -1) { outputStream.write(buffer, 0, length); } inputStream.close(); outputStream.close(); System.out.println("Image downloaded successfully!"); } catch (IOException e) { e.printStackTrace(); } }
解决方法
1. 获取图片实际URL
很多网站会用相对路径、动态跳转链接或懒加载属性存储图片地址,可按以下方式处理:
- 补全相对路径:如果
src是相对路径,拼接网站基础URL形成完整链接String baseUrl = "https://目标网站域名.com"; // 替换为实际网站域名 String imageUrl = image.getAttribute("src"); if (!imageUrl.startsWith("http")) { imageUrl = baseUrl + imageUrl; } - 读取懒加载属性:部分网站会把真实图片地址存在
data-src、data-original等属性中String imageUrl = image.getAttribute("data-src"); if (imageUrl == null || imageUrl.isEmpty()) { imageUrl = image.getAttribute("src"); }
2. 直接截取img元素截图
如果无法获取正确的图片URL,可利用Selenium的截图API直接截取元素可视化内容:
import java.nio.file.Files; import java.nio.file.StandardCopyOption; WebElement image = webDriver.findElement(By.xpath("//img")); File screenshot = image.getScreenshotAs(OutputType.FILE); File destination = new File(directory + "downloadedImage.png"); Files.copy(screenshot.toPath(), destination.toPath(), StandardCopyOption.REPLACE_EXISTING);
3. 模拟浏览器请求头与Cookie
部分网站会通过Cookie、请求头做防盗链或身份验证,直接用URLConnection请求会返回HTML页面。此时需复制浏览器的Cookie和请求头:
URL url = new URL(imageUrl); HttpURLConnection connection = (HttpURLConnection) url.openConnection(); // 复制浏览器Cookie String cookieString = ""; Set<Cookie> cookies = webDriver.manage().getCookies(); for (Cookie cookie : cookies) { cookieString += cookie.getName() + "=" + cookie.getValue() + "; "; } connection.setRequestProperty("Cookie", cookieString); // 添加浏览器标识与来源页 connection.setRequestProperty("User-Agent", webDriver.executeScript("return navigator.userAgent").toString()); connection.setRequestProperty("Referer", webDriver.getCurrentUrl()); // 后续读取流代码不变 InputStream inputStream = connection.getInputStream(); // ...
内容的提问来源于stack exchange,提问作者javabeginer
相关产品推荐
相关产品推荐

