You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:使用BS4提取Offspring网站商品信息遇难题

Solution for Scraping Offspring Product Details with BeautifulSoup

Hey there! I’ve spent a lot of time scraping e-commerce sites with BS4, so let’s fix your issues with extracting product titles and images from that Offspring page.

First, Let’s Break Down the Common Pitfalls You’re Facing

  1. Image extraction failing: Most modern e-commerce sites use lazy loading for images, meaning the actual image URL is stored in a data-* attribute (like data-src) instead of the standard src (which often points to a placeholder).
  2. Missing product titles: You need to target the specific element that holds the title within each product’s container—don’t try to grab it globally.

Step-by-Step Fixes & Full Code Example

Here’s a robust script that grabs all the details you need, with explanations for each part:

import requests
from bs4 import BeautifulSoup

# Target URL with headers to avoid being blocked
base_url = "https://www.offspring.co.uk/view/category/offspring_catalog/NEWIN?pageSize=100&sort=-releasedate"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

# Fetch and parse the page
response = requests.get(base_url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

# Locate every product tile on the page (this is your core container)
product_tiles = soup.find_all("article", class_="product-tile")

for tile in product_tiles:
    # 1. Extract Product URL (you had this right, but wrapping it in the tile makes it safer)
    product_url = tile.find("a").attrs["href"]
    
    # 2. Extract Price (your original code works here, just tied to the tile)
    price = tile.find("div", class_="price js-price").text.strip()
    
    # 3. Extract Product Title
    # The title is stored in the `title` attribute of the main product link
    product_link = tile.find("a", class_="product-link")
    title = product_link.get("title").strip()
    # If the title attribute is missing, fall back to the text inside the link:
    # title = product_link.find("span", class_="product-title").text.strip()
    
    # 4. Extract Image URL (fix for lazy loading)
    img_tag = tile.find("img")
    # Check for lazy load attributes first, then fall back to src
    image_url = img_tag.get("data-src") or img_tag.get("src")
    
    # Print or store the results
    print(f"Product Title: {title}")
    print(f"Price: {price}")
    print(f"Image URL: {image_url}")
    print(f"Product URL: {product_url}\n")

Key Explanations

  • Product Container: By targeting <article class="product-tile">, you ensure you’re only pulling data from one product at a time—no cross-contamination between items.
  • Title Extraction: The "Jordan Zoom 92 Trainer" title is embedded in the title attribute of the product’s main link. If that ever changes, you can switch to grabbing the text from the <span class="product-title"> inside the link.
  • Image Fix: The site uses data-src for lazy loading, so we check that attribute first. If it’s not present (for already-loaded images), we fall back to src.

Quick Notes

  • Always include a User-Agent header to mimic a real browser—this helps avoid being blocked by the site’s anti-scraping measures.
  • If the page structure changes (which e-commerce sites often do), use your browser’s dev tools (right-click → Inspect) to recheck the class names and attributes.

内容的提问来源于stack exchange,提问作者fdezgar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 23:57:45