求助:使用BS4提取Offspring网站商品信息遇难题
Solution for Scraping Offspring Product Details with BeautifulSoup
Hey there! I’ve spent a lot of time scraping e-commerce sites with BS4, so let’s fix your issues with extracting product titles and images from that Offspring page.
First, Let’s Break Down the Common Pitfalls You’re Facing
- Image extraction failing: Most modern e-commerce sites use lazy loading for images, meaning the actual image URL is stored in a
data-*attribute (likedata-src) instead of the standardsrc(which often points to a placeholder). - Missing product titles: You need to target the specific element that holds the title within each product’s container—don’t try to grab it globally.
Step-by-Step Fixes & Full Code Example
Here’s a robust script that grabs all the details you need, with explanations for each part:
import requests from bs4 import BeautifulSoup # Target URL with headers to avoid being blocked base_url = "https://www.offspring.co.uk/view/category/offspring_catalog/NEWIN?pageSize=100&sort=-releasedate" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # Fetch and parse the page response = requests.get(base_url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # Locate every product tile on the page (this is your core container) product_tiles = soup.find_all("article", class_="product-tile") for tile in product_tiles: # 1. Extract Product URL (you had this right, but wrapping it in the tile makes it safer) product_url = tile.find("a").attrs["href"] # 2. Extract Price (your original code works here, just tied to the tile) price = tile.find("div", class_="price js-price").text.strip() # 3. Extract Product Title # The title is stored in the `title` attribute of the main product link product_link = tile.find("a", class_="product-link") title = product_link.get("title").strip() # If the title attribute is missing, fall back to the text inside the link: # title = product_link.find("span", class_="product-title").text.strip() # 4. Extract Image URL (fix for lazy loading) img_tag = tile.find("img") # Check for lazy load attributes first, then fall back to src image_url = img_tag.get("data-src") or img_tag.get("src") # Print or store the results print(f"Product Title: {title}") print(f"Price: {price}") print(f"Image URL: {image_url}") print(f"Product URL: {product_url}\n")
Key Explanations
- Product Container: By targeting
<article class="product-tile">, you ensure you’re only pulling data from one product at a time—no cross-contamination between items. - Title Extraction: The "Jordan Zoom 92 Trainer" title is embedded in the
titleattribute of the product’s main link. If that ever changes, you can switch to grabbing the text from the<span class="product-title">inside the link. - Image Fix: The site uses
data-srcfor lazy loading, so we check that attribute first. If it’s not present (for already-loaded images), we fall back tosrc.
Quick Notes
- Always include a
User-Agentheader to mimic a real browser—this helps avoid being blocked by the site’s anti-scraping measures. - If the page structure changes (which e-commerce sites often do), use your browser’s dev tools (right-click → Inspect) to recheck the class names and attributes.
内容的提问来源于stack exchange,提问作者fdezgar
相关产品推荐
相关产品推荐

