You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium或其他爬虫技术保存div背景图片为文件?

How to Save a div's Background Image with Selenium (Even for One-Time URLs)

Hey there! I’ve dealt with exactly this kind of one-time image URL scenario before—super frustrating when you can’t just right-click and save, but the image is clearly loaded in the browser. Let’s go through two reliable methods to grab that background image using Selenium and browser tools.

Method 1: Use JavaScript + Canvas to Extract the Image (Selenium)

Since the browser already has the image cached once the page loads, we can use JavaScript to draw it onto a canvas and export it as a Base64 string, then decode and save it locally.

Here’s a Python example:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import base64
import os

# Launch Chrome and load your page
driver = webdriver.Chrome()
driver.get("YOUR_PAGE_URL_HERE")

# Wait for the target div to load completely
wait = WebDriverWait(driver, 15)
book_div = wait.until(EC.presence_of_element_located((By.ID, "BookImage")))

# Run JavaScript to get the background image as Base64
extract_image_js = """
return new Promise((resolve, reject) => {
    const div = document.getElementById('BookImage');
    // Extract the URL from the background-image style
    const bgUrl = div.style.backgroundImage.replace(/url\\(["']?([^"']*)["']?\\)/, '$1');
    
    const img = new Image();
    // Handle cross-origin issues if needed
    img.crossOrigin = 'anonymous';
    
    img.onload = () => {
        const canvas = document.createElement('canvas');
        canvas.width = img.width;
        canvas.height = img.height;
        const ctx = canvas.getContext('2d');
        ctx.drawImage(img, 0, 0);
        // Export as PNG Base64
        resolve(canvas.toDataURL('image/png'));
    };
    
    img.onerror = () => reject('Failed to load background image');
    img.src = bgUrl;
});
"""

base64_img = driver.execute_script(extract_image_js)

# Decode Base64 and save the image
image_data = base64.b64decode(base64_img.split(',')[1])
save_path = os.path.join(os.getcwd(), 'book_chapter_image.png')

with open(save_path, 'wb') as file:
    file.write(image_data)

print(f"Image saved to: {save_path}")
driver.quit()

Notes for Method 1:

  • If the image is hosted on a different domain, you might hit CORS errors unless the server allows cross-origin requests. If that happens, try the second method instead.
  • Make sure the page is fully loaded before running the script—adjust the wait time if the image loads slowly.

Method 2: Use Chrome DevTools Protocol (CDP) to Grab Cached Resources

Since you can see the image in the Sources tab, that means Chrome has already downloaded it. We can use CDP (built into Selenium for Chrome) to fetch the raw image data directly from the browser’s network cache.

Here’s how to do it in Python:

from selenium import webdriver
import base64
import os

# Initialize Chrome with CDP enabled
options = webdriver.ChromeOptions()
driver = webdriver.Chrome(options=options)

# Enable network monitoring via CDP
driver.execute_cdp_cmd('Network.enable', {})

# Load your page
driver.get("YOUR_PAGE_URL_HERE")

# Wait a bit for all resources to load (adjust as needed)
driver.implicitly_wait(10)

# Fetch all network requests made by the page
all_requests = driver.execute_cdp_cmd('Network.getRequests', {})['requests']

# Find the request matching your background image URL (filter by your URL pattern)
target_request = None
for req in all_requests:
    if 'book_chapter_image' in req['url']:
        target_request = req
        break

if target_request:
    # Get the raw response body of the image
    response = driver.execute_cdp_cmd('Network.getResponseBody', {
        'requestId': target_request['requestId']
    })
    
    # Decode and save the image
    if response['base64Encoded']:
        image_data = base64.b64decode(response['body'])
    else:
        image_data = response['body'].encode('utf-8')
    
    save_path = os.path.join(os.getcwd(), 'book_chapter_image_cdp.png')
    with open(save_path, 'wb') as file:
        file.write(image_data)
    
    print(f"Image saved via CDP to: {save_path}")
else:
    print("Couldn't find the target image request in network logs")

driver.quit()

Why This Works:

CDP lets us access the browser’s internal network logs and cached resources, so we don’t have to rely on the one-time URL anymore. This method avoids CORS issues entirely because we’re pulling the image directly from what Chrome already loaded.

Final Tips:

  • If the image loads dynamically (e.g., after scrolling or a button click), make sure to trigger that action in Selenium before running the extraction code.
  • Test with a short wait time first, but adjust if the image takes longer to load—you don’t want to run the script before the image is cached.

内容的提问来源于stack exchange,提问作者BookSword

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:00:45