You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scraping Scryfall图片URL为空求助:滚动加载致ZNR集保存失败

Fixing Lazy-Loaded Image Scraping for Scryfall Sets (Zendikar Rising Issue)

Hey there! I see exactly what's going on here. Your script works fine for smaller sets, but Zendikar Rising has way more cards, and Scryfall uses JavaScript-based lazy loading for images—meaning images only load when they scroll into the browser's viewport. When you use requests.get() to fetch the page, you're only getting the initial HTML with placeholder image URLs (or empty ones) for cards that aren't in the initial view. That's why you're seeing those "Failed to save image" errors and missing cards.

Let's break down how to fix this:

Step 1: Use a Tool That Handles JavaScript Rendering

Instead of requests, we'll use Selenium—it simulates a real browser, so we can scroll the page to trigger all lazy-loaded images to load. You'll need to install Selenium and download a browser driver (like ChromeDriver for Chrome).

First, install the package:

pip install selenium

Step 2: Modified Script with Selenium

Here's the updated code that handles lazy loading, ensures card names and URLs stay paired, and has better error handling:

import requests
import os
from os.path import basename
from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

path = os.getcwd()
print(f"The current working directory is {path}")
url = 'https://scryfall.com/sets'

# Set up Selenium (headless mode to avoid opening a browser window)
options = webdriver.ChromeOptions()
options.add_argument('--headless=new')
options.add_argument('--disable-gpu')
driver = webdriver.Chrome(options=options)

try:
    # Fetch the sets page
    driver.get(url)
    WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.TAG_NAME, 'a')))
    
    # Gather all set URLs
    soup = BeautifulSoup(driver.page_source, 'html.parser')
    Urls = []
    for link in soup.findAll('a'):
        href = link.get('href')
        if href and 'https://scryfall.com/sets/' in href and href not in Urls:
            Urls.append(href)

    # Loop through each set
    for Url in Urls:
        driver.get(Url)
        # Wait for set header to load
        WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.CLASS_NAME, 'set-header-title-h1')))
        
        # Scroll the page multiple times to trigger all lazy-loaded images
        last_height = driver.execute_script("return document.body.scrollHeight")
        while True:
            # Scroll to bottom
            driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
            # Wait for images to load
            time.sleep(2)
            # Check if scroll height changed
            new_height = driver.execute_script("return document.body.scrollHeight")
            if new_height == last_height:
                break
            last_height = new_height

        # Now parse the fully loaded page
        soup = BeautifulSoup(driver.page_source, 'html.parser')
        
        # Create set directory
        set_title = soup.find('h1', {'class': 'set-header-title-h1'}).get_text(strip=True)
        set_title = set_title.replace(':', '').replace(' ', '')
        set_dir = f"{path}\\{set_title}"
        try:
            os.mkdir(set_dir)
            print(f"Successfully created the directory {set_dir}")
        except OSError:
            print(f"Directory {set_dir} already exists, skipping creation")

        # Gather card names and image URLs (note: Scryfall uses data-src for lazy-loaded images)
        card_elements = soup.find_all('img', {'class': 'card-image'})
        for card in card_elements:
            card_name = card.get('alt')
            # Get the real image URL from data-src (fallback to src if needed)
            img_url = card.get('data-src') or card.get('src')
            
            if not card_name or not img_url:
                print(f"Skipping card with missing name or URL: {card_name} | {img_url}")
                continue
            
            # Save the image
            fn = f"{card_name}.png"
            try:
                img_data = requests.get(img_url).content
                with open(f"{set_dir}\\{basename(fn)}", "wb") as f:
                    f.write(img_data)
                # Uncomment to log successful saves
                # print(f"Saved {fn} to {set_dir}")
            except Exception as e:
                print(f"Failed to save image {fn} from url {img_url}: {str(e)}")

    print("Completed With No Errors")

finally:
    # Make sure to close the browser
    driver.quit()

Key Improvements Explained

  • Selenium with Headless Chrome: Simulates browsing, scrolls the page to load all lazy images without popping up a visible browser window.
  • Scroll Loop: Keeps scrolling until the page height stops changing, ensuring every card's image has been triggered to load.
  • Data-Src Handling: Scryfall stores the real image URL in data-src for lazy-loaded images, so we prioritize that over the placeholder src value.
  • Granular Error Handling: Skips individual problematic cards instead of exiting the entire script, so you don't lose progress on the whole set.
  • Wait Conditions: Uses WebDriverWait to ensure critical page elements are loaded before parsing, avoiding race conditions where the script tries to read content that hasn't rendered yet.

Important Notes

  • You'll need to download the ChromeDriver that matches your Chrome version, and make sure it's in your system PATH or specify its path directly in the webdriver.Chrome() call (e.g., webdriver.Chrome(executable_path='path/to/chromedriver', options=options)).
  • Adjust the time.sleep(2) value if needed—slower internet connections might require a longer wait for images to fully load after each scroll.
  • Large sets like Zendikar Rising will take a minute or two to process, but this script will capture every card in the set.

内容的提问来源于stack exchange,提问作者Sitruc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 08:12:46