Python爬取斯凯奇网站无法完整打印商品名称、图片及价格问题求助
Fixing Your Skechers Web Scraping Code
Let’s walk through the issues in your code and fix them to successfully pull product names, images, and prices from the page:
Key Problems in Your Original Code
- Multi-class selector error:
find_elements_by_class_namecan’t handle multiple class names separated by spaces. You need to use a CSS selector instead. - Loop variable typo: You named the loop variable
vitbut usedvideoinside the loop—this would throw an undefined variable error. - Global element search: Using
//in XPath searches the entire page, so you were always grabbing the first product’s details instead of the current tile’s. - Incorrect image extraction: You tried to get the image’s text, but the actual image URL is stored in the
srcattribute. - Deprecated Selenium methods: Older
find_elements_by_*methods are outdated; use theByclass for selectors instead. - Missing page load wait: The page might not fully render before you try to extract elements, leading to empty results.
Corrected Code
import time from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.service import Service # Required for Selenium 4+ # Set up Chrome driver (adjust path to your chromedriver.exe) service = Service('D:/chromedriver.exe') driver = webdriver.Chrome(service=service) # For older Selenium versions, use: driver = webdriver.Chrome('D:/chromedriver.exe') url = 'https://www.skechers.com/women/shoes/athletic-sneakers/?start=0&sz=168' driver.get(url) # Wait for the page to fully load dynamic content (adjust time if needed) time.sleep(5) # Find all product tiles using CSS selector (supports multiple classes) product_tiles = driver.find_elements(By.CSS_SELECTOR, 'div.col-6.col-sm-4.col-xl-3.mb-2.mb-md-1.mb-lg-4.px-lg-3') print(f"Found {len(product_tiles)} products:\n") for tile in product_tiles: try: # Extract title from the tile-specific link title = tile.find_element(By.CSS_SELECTOR, 'a.link.c-product-tile__title').text.strip() # Get image URL from the img tag's src attribute image_url = tile.find_element(By.CSS_SELECTOR, 'img.tile-image.c-product-tile__img').get_attribute('src') # Extract price from the tile's value span price = tile.find_element(By.CSS_SELECTOR, 'span.value').text.strip() # Print formatted product details print(f"Title: {title}") print(f"Image URL: {image_url}") print(f"Price: {price}\n") except Exception as e: print(f"Skipping product due to error: {str(e)}\n") continue # Clean up: close the browser driver.quit()
How This Works
- CSS Selectors: We use
By.CSS_SELECTORto target elements with multiple classes, which is more reliable thanclass_namefor multi-class elements. - Tile-Specific Search: By calling
tile.find_element()instead ofdriver.find_element(), we only search within the current product tile—ensuring we get the right details for each item. - Image URL Extraction: Instead of
.text, we useget_attribute('src')to retrieve the actual image link. - Error Handling: The
try-exceptblock ensures that if one product has missing elements, the loop continues to process the rest of the items. - Page Load Wait:
time.sleep(5)gives the page enough time to render all dynamic content (adjust this based on your internet speed).
Pro Tip for Better Reliability
Replace the time.sleep() with explicit waits to avoid unnecessary delays and ensure elements are ready before extraction:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Wait up to 10 seconds for product tiles to load WebDriverWait(driver, 10).until( EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'div.col-6.col-sm-4.col-xl-3.mb-2.mb-md-1.mb-lg-4.px-lg-3')) )
内容的提问来源于stack exchange,提问作者asad khan
相关产品推荐
相关产品推荐

