如何用Selenium或其他爬虫技术保存div背景图片为文件?
Hey there! I’ve dealt with exactly this kind of one-time image URL scenario before—super frustrating when you can’t just right-click and save, but the image is clearly loaded in the browser. Let’s go through two reliable methods to grab that background image using Selenium and browser tools.
Method 1: Use JavaScript + Canvas to Extract the Image (Selenium)
Since the browser already has the image cached once the page loads, we can use JavaScript to draw it onto a canvas and export it as a Base64 string, then decode and save it locally.
Here’s a Python example:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import base64 import os # Launch Chrome and load your page driver = webdriver.Chrome() driver.get("YOUR_PAGE_URL_HERE") # Wait for the target div to load completely wait = WebDriverWait(driver, 15) book_div = wait.until(EC.presence_of_element_located((By.ID, "BookImage"))) # Run JavaScript to get the background image as Base64 extract_image_js = """ return new Promise((resolve, reject) => { const div = document.getElementById('BookImage'); // Extract the URL from the background-image style const bgUrl = div.style.backgroundImage.replace(/url\\(["']?([^"']*)["']?\\)/, '$1'); const img = new Image(); // Handle cross-origin issues if needed img.crossOrigin = 'anonymous'; img.onload = () => { const canvas = document.createElement('canvas'); canvas.width = img.width; canvas.height = img.height; const ctx = canvas.getContext('2d'); ctx.drawImage(img, 0, 0); // Export as PNG Base64 resolve(canvas.toDataURL('image/png')); }; img.onerror = () => reject('Failed to load background image'); img.src = bgUrl; }); """ base64_img = driver.execute_script(extract_image_js) # Decode Base64 and save the image image_data = base64.b64decode(base64_img.split(',')[1]) save_path = os.path.join(os.getcwd(), 'book_chapter_image.png') with open(save_path, 'wb') as file: file.write(image_data) print(f"Image saved to: {save_path}") driver.quit()
Notes for Method 1:
- If the image is hosted on a different domain, you might hit CORS errors unless the server allows cross-origin requests. If that happens, try the second method instead.
- Make sure the page is fully loaded before running the script—adjust the wait time if the image loads slowly.
Method 2: Use Chrome DevTools Protocol (CDP) to Grab Cached Resources
Since you can see the image in the Sources tab, that means Chrome has already downloaded it. We can use CDP (built into Selenium for Chrome) to fetch the raw image data directly from the browser’s network cache.
Here’s how to do it in Python:
from selenium import webdriver import base64 import os # Initialize Chrome with CDP enabled options = webdriver.ChromeOptions() driver = webdriver.Chrome(options=options) # Enable network monitoring via CDP driver.execute_cdp_cmd('Network.enable', {}) # Load your page driver.get("YOUR_PAGE_URL_HERE") # Wait a bit for all resources to load (adjust as needed) driver.implicitly_wait(10) # Fetch all network requests made by the page all_requests = driver.execute_cdp_cmd('Network.getRequests', {})['requests'] # Find the request matching your background image URL (filter by your URL pattern) target_request = None for req in all_requests: if 'book_chapter_image' in req['url']: target_request = req break if target_request: # Get the raw response body of the image response = driver.execute_cdp_cmd('Network.getResponseBody', { 'requestId': target_request['requestId'] }) # Decode and save the image if response['base64Encoded']: image_data = base64.b64decode(response['body']) else: image_data = response['body'].encode('utf-8') save_path = os.path.join(os.getcwd(), 'book_chapter_image_cdp.png') with open(save_path, 'wb') as file: file.write(image_data) print(f"Image saved via CDP to: {save_path}") else: print("Couldn't find the target image request in network logs") driver.quit()
Why This Works:
CDP lets us access the browser’s internal network logs and cached resources, so we don’t have to rely on the one-time URL anymore. This method avoids CORS issues entirely because we’re pulling the image directly from what Chrome already loaded.
Final Tips:
- If the image loads dynamically (e.g., after scrolling or a button click), make sure to trigger that action in Selenium before running the extraction code.
- Test with a short wait time first, but adjust if the image takes longer to load—you don’t want to run the script before the image is cached.
内容的提问来源于stack exchange,提问作者BookSword

