You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup解析亚马逊商品参数(品牌/尺寸等)失败求助

解决亚马逊商品参数提取问题

Hey there! I see you've got the basics down with parsing Amazon's search results, but hitting a snag pulling those specific product specs like Brand, Display Size, etc. Let's fix that together.

问题分析

The issue is that your current code only fetches and parses the page, but doesn't target the individual product containers or the spec blocks within them. Each search result item has its own s-product-specs-view section, so we need to loop through each product first, then extract specs from the relevant block inside each one.

完整解决方案代码

Here's an updated version of your code that will extract all the data you need, including the missing specs:

from bs4 import BeautifulSoup
from urllib.request import Request, urlopen

site = 'https://www.amazon.com/s?k=Apple&rh=n%3A2407749011'
hdr = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'}
req = Request(site, headers=hdr)
page = urlopen(req)
soup = BeautifulSoup(page, 'html.parser')

# List to store all product data
products = []

# Loop through each product in the search results
for product in soup.find_all('div', {'data-component-type': 's-search-result'}):
    product_data = {}
    
    # Extract product name (you probably already have this)
    name_elem = product.find('span', class_='a-size-medium a-color-base a-text-normal')
    product_data['Name'] = name_elem.get_text(strip=True) if name_elem else 'N/A'
    
    # Extract price (your existing logic)
    price_elem = product.find('span', class_='a-price-whole')
    product_data['Price'] = price_elem.get_text(strip=True) if price_elem else 'N/A'
    
    # Extract review count (your existing logic)
    review_elem = product.find('span', class_='a-size-base s-underline-text')
    product_data['Review Count'] = review_elem.get_text(strip=True) if review_elem else 'N/A'
    
    # Now extract the specs: Brand, Display Size, etc.
    spec_block = product.find('div', class_='s-product-specs-view')
    if spec_block:
        # Get all spec items in the block
        spec_items = spec_block.find_all('div', class_='sg-col-inner')
        for item in spec_items:
            # Get the label (e.g., Brand) and value (e.g., Apple)
            label = item.find('span', dir='auto').get_text(strip=True)
            value = item.find('span', class_='a-text-bold').get_text(strip=True)
            product_data[label] = value
    else:
        # If no spec block, mark specs as N/A
        product_data['Brand'] = 'N/A'
        product_data['Display Size'] = 'N/A'
        product_data['Capacity'] = 'N/A'
        product_data['Color'] = 'N/A'
    
    products.append(product_data)

# Print out the results to verify
for idx, prod in enumerate(products, 1):
    print(f"Product {idx}:")
    for key, value in prod.items():
        print(f"  {key}: {value}")
    print("-"*50)

关键步骤解释

  • Targeting individual products: We use soup.find_all('div', {'data-component-type': 's-search-result'}) to get every product card in the search results. This ensures we're extracting data per-item, not from random parts of the page.
  • Accessing the spec block: For each product, we look for the s-product-specs-view div. If it exists, we loop through each sg-col-inner div inside it.
  • Extracting spec labels and values: Each spec item has a regular span for the label (like "Brand") and a bold span for the value. We grab both and add them to our product dictionary.
  • Handling missing data: We add checks for cases where a product might not have a spec block or specific elements (using if elem else 'N/A') to avoid errors.

注意事项

  • Amazon's anti-scraping measures: Amazon actively blocks scrapers, so you might need to rotate your User-Agent header or add delays between requests if you're making multiple calls. Avoid hitting the site too frequently.
  • Page structure changes: Amazon updates its HTML structure regularly, so if this code stops working later, you'll need to inspect the page again and adjust the selectors.

内容的提问来源于stack exchange,提问作者KHW

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:35:06