非Headless模式下Selenium无法截取网页完整截图的解决方案求助
解决方案:非Headless模式下截取完整网页截图
嘿,我完全懂你这个困扰!非Headless模式下被显示器分辨率卡着,没法把窗口拉到足够高来截完整页,之前那些Headless的方法确实没法直接套用。我之前也踩过这个坑,给你分享两个亲测有效的解决方案:
方案1:使用Chrome DevTools Protocol(CDP)直接截取全页
这个方法是目前最省心的,它绕过了窗口尺寸的限制,直接从浏览器的渲染层面捕获整个页面,不管你是不是用Headless模式都能生效。
import os import time from selenium import webdriver from selenium.webdriver.chrome.service import Service # 初始化非Headless浏览器(--start-maximized可选,不影响截图结果) chrome_options = webdriver.ChromeOptions() chrome_options.add_argument("--start-maximized") chromedriver_path = os.path.join(os.path.dirname(os.path.abspath(__file__)), 'chromedriver.exe') driver = webdriver.Chrome(service=Service(chromedriver_path), options=chrome_options) url = 'https://stackoverflow.com/' driver.get(url) time.sleep(2) # 等待页面完全加载,可根据实际情况调整 # 调用CDP的Page.captureScreenshot方法,开启视口外捕获 screenshot_data = driver.execute_cdp_cmd( "Page.captureScreenshot", {"captureBeyondViewport": True, "fromSurface": True} ) # 将返回的十六进制数据转为字节并保存 with open("full_page_screenshot.png", "wb") as file: file.write(bytes.fromhex(screenshot_data['data'])) driver.quit()
为什么这个方法有效?
Chrome DevTools Protocol(CDP)是浏览器提供的一套调试接口,Page.captureScreenshot里的captureBeyondViewport参数专门用来指示浏览器捕获当前视口之外的页面内容,完全不受显示器分辨率或当前窗口大小的限制,非Headless模式下完美适配。
方案2:滚动页面分区域截取后拼接
如果因为某些原因无法使用CDP,这个变通方法也能解决问题——通过多次滚动页面,截取每个可见区域的截图,最后用图片处理工具拼接成完整页面。
import os import time from selenium import webdriver from selenium.webdriver.chrome.service import Service from PIL import Image # 初始化非Headless浏览器 chrome_options = webdriver.ChromeOptions() chrome_options.add_argument("--start-maximized") chromedriver_path = os.path.join(os.path.dirname(os.path.abspath(__file__)), 'chromedriver.exe') driver = webdriver.Chrome(service=Service(chromedriver_path), options=chrome_options) url = 'https://stackoverflow.com/' driver.get(url) time.sleep(2) # 等待页面加载 # 获取页面总高度和当前窗口可见高度 total_page_height = driver.execute_script("return document.body.scrollHeight") window_visible_height = driver.execute_script("return window.innerHeight") current_scroll_pos = 0 temp_screenshots = [] # 循环滚动并截取每个区域 while current_scroll_pos < total_page_height: # 保存当前可见区域的截图 temp_path = f"temp_screenshot_{current_scroll_pos}.png" driver.save_screenshot(temp_path) temp_screenshots.append(Image.open(temp_path)) # 滚动到下一个区域 current_scroll_pos += window_visible_height driver.execute_script(f"window.scrollTo(0, {current_scroll_pos});") time.sleep(0.5) # 等待滚动后内容加载,避免截图空白 # 拼接所有临时截图 if temp_screenshots: full_width = temp_screenshots[0].width full_height = sum(img.height for img in temp_screenshots) full_page_image = Image.new('RGB', (full_width, full_height)) current_y = 0 for img in temp_screenshots: full_page_image.paste(img, (0, current_y)) current_y += img.height full_page_image.save("full_page_combined.png") # 清理临时文件 for temp_path in [f"temp_screenshot_{i}.png" for i in range(0, current_scroll_pos, window_visible_height)]: if os.path.exists(temp_path): os.remove(temp_path) driver.quit()
注意事项
- 如果页面有固定定位的导航栏/工具栏,滚动时会重复出现在每个截图里,需要计算导航栏高度,调整每次滚动的距离(比如
current_scroll_pos += window_visible_height - navbar_height) - 滚动后的等待时间
time.sleep(0.5)可根据页面加载速度调整,避免截图出现空白
内容的提问来源于stack exchange,提问作者StyleZ
相关产品推荐
相关产品推荐

