Scraping Scryfall图片URL为空求助:滚动加载致ZNR集保存失败
Hey there! I see exactly what's going on here. Your script works fine for smaller sets, but Zendikar Rising has way more cards, and Scryfall uses JavaScript-based lazy loading for images—meaning images only load when they scroll into the browser's viewport. When you use requests.get() to fetch the page, you're only getting the initial HTML with placeholder image URLs (or empty ones) for cards that aren't in the initial view. That's why you're seeing those "Failed to save image" errors and missing cards.
Let's break down how to fix this:
Step 1: Use a Tool That Handles JavaScript Rendering
Instead of requests, we'll use Selenium—it simulates a real browser, so we can scroll the page to trigger all lazy-loaded images to load. You'll need to install Selenium and download a browser driver (like ChromeDriver for Chrome).
First, install the package:
pip install selenium
Step 2: Modified Script with Selenium
Here's the updated code that handles lazy loading, ensures card names and URLs stay paired, and has better error handling:
import requests import os from os.path import basename from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time path = os.getcwd() print(f"The current working directory is {path}") url = 'https://scryfall.com/sets' # Set up Selenium (headless mode to avoid opening a browser window) options = webdriver.ChromeOptions() options.add_argument('--headless=new') options.add_argument('--disable-gpu') driver = webdriver.Chrome(options=options) try: # Fetch the sets page driver.get(url) WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.TAG_NAME, 'a'))) # Gather all set URLs soup = BeautifulSoup(driver.page_source, 'html.parser') Urls = [] for link in soup.findAll('a'): href = link.get('href') if href and 'https://scryfall.com/sets/' in href and href not in Urls: Urls.append(href) # Loop through each set for Url in Urls: driver.get(Url) # Wait for set header to load WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.CLASS_NAME, 'set-header-title-h1'))) # Scroll the page multiple times to trigger all lazy-loaded images last_height = driver.execute_script("return document.body.scrollHeight") while True: # Scroll to bottom driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # Wait for images to load time.sleep(2) # Check if scroll height changed new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # Now parse the fully loaded page soup = BeautifulSoup(driver.page_source, 'html.parser') # Create set directory set_title = soup.find('h1', {'class': 'set-header-title-h1'}).get_text(strip=True) set_title = set_title.replace(':', '').replace(' ', '') set_dir = f"{path}\\{set_title}" try: os.mkdir(set_dir) print(f"Successfully created the directory {set_dir}") except OSError: print(f"Directory {set_dir} already exists, skipping creation") # Gather card names and image URLs (note: Scryfall uses data-src for lazy-loaded images) card_elements = soup.find_all('img', {'class': 'card-image'}) for card in card_elements: card_name = card.get('alt') # Get the real image URL from data-src (fallback to src if needed) img_url = card.get('data-src') or card.get('src') if not card_name or not img_url: print(f"Skipping card with missing name or URL: {card_name} | {img_url}") continue # Save the image fn = f"{card_name}.png" try: img_data = requests.get(img_url).content with open(f"{set_dir}\\{basename(fn)}", "wb") as f: f.write(img_data) # Uncomment to log successful saves # print(f"Saved {fn} to {set_dir}") except Exception as e: print(f"Failed to save image {fn} from url {img_url}: {str(e)}") print("Completed With No Errors") finally: # Make sure to close the browser driver.quit()
Key Improvements Explained
- Selenium with Headless Chrome: Simulates browsing, scrolls the page to load all lazy images without popping up a visible browser window.
- Scroll Loop: Keeps scrolling until the page height stops changing, ensuring every card's image has been triggered to load.
- Data-Src Handling: Scryfall stores the real image URL in
data-srcfor lazy-loaded images, so we prioritize that over the placeholdersrcvalue. - Granular Error Handling: Skips individual problematic cards instead of exiting the entire script, so you don't lose progress on the whole set.
- Wait Conditions: Uses
WebDriverWaitto ensure critical page elements are loaded before parsing, avoiding race conditions where the script tries to read content that hasn't rendered yet.
Important Notes
- You'll need to download the ChromeDriver that matches your Chrome version, and make sure it's in your system PATH or specify its path directly in the
webdriver.Chrome()call (e.g.,webdriver.Chrome(executable_path='path/to/chromedriver', options=options)). - Adjust the
time.sleep(2)value if needed—slower internet connections might require a longer wait for images to fully load after each scroll. - Large sets like Zendikar Rising will take a minute or two to process, but this script will capture every card in the set.
内容的提问来源于stack exchange,提问作者Sitruc

