You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升Streamlit程序运行速度及优化亚马逊商品爬虫效果?

亚马逊商品名称爬虫优化方案

一、解决「获取数量少」的问题

1. 完善请求头,模拟真实浏览器

仅用User-Agent容易被亚马逊反爬机制识别,补充更多浏览器请求头字段,降低拦截概率:

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',
    'Accept-Language': 'en-US,en;q=0.5',
    'Referer': 'https://www.amazon.com/',
    'Upgrade-Insecure-Requests': '1',
    'Connection': 'keep-alive'
}

2. 增加反爬应对机制

  • 随机延迟:每次请求后随机等待1-3秒,避免请求频率触发限制
  • 重试机制:请求失败或未获取标题时,重试1-2次,重试时可更换请求头或代理
  • 代理IP轮换:针对大量ASIN爬取,用代理池切换IP,避免单IP被封禁

示例整合代码:

import time
import random
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

def create_session():
    session = requests.Session()
    # 配置重试策略
    retry = Retry(total=2, backoff_factor=1, status_forcelist=[429, 500, 502, 503, 504])
    adapter = HTTPAdapter(max_retries=retry)
    session.mount('https://', adapter)
    return session

def scrape_product_names(asins, csv_path='product_names.csv'):
    product_mapping = {}
    session = create_session()
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',
        'Accept-Language': 'en-US,en;q=0.5',
        'Referer': 'https://www.amazon.com/',
        'Upgrade-Insecure-Requests': '1',
        'Connection': 'keep-alive'
    }
    for asin in asins:
        try:
            url = f"https://www.amazon.com/dp/{asin}"
            time.sleep(random.uniform(1,3))  # 随机延迟
            response = session.get(url, headers=headers)
            response.raise_for_status()  # 抛出HTTP错误
            soup = BeautifulSoup(response.content, 'html.parser')
            # 优先尝试原选择器
            product_title = soup.find("span", {"id": "productTitle"})
            if product_title:
                product_name = product_title.get_text().strip().split('|')[0].strip()
                product_mapping[asin] = product_name
            else:
                # 备选选择器适配不同页面结构
                alt_title = soup.find("h1", class_="a-size-large a-spacing-none")
                product_mapping[asin] = alt_title.get_text().strip() if alt_title else asin
        except Exception as e:
            print(f"处理ASIN {asin}出错: {str(e)}")
            product_mapping[asin] = asin
    return product_mapping

3. 适配多页面结构

部分商品页面(如第三方卖家页)的标题可能不在productTitle节点,增加备选选择器,比如span.a-text-title、h1.a-size-medium等,覆盖更多页面场景。

二、解决「运行耗时久」的问题

1. 改用异步请求

用aiohttp替代requests实现并发请求,大幅缩短大量ASIN的爬取总耗时:

import aiohttp
import asyncio
from bs4 import BeautifulSoup

async def fetch_asin(session, asin, headers):
    url = f"https://www.amazon.com/dp/{asin}"
    try:
        await asyncio.sleep(random.uniform(0.5,2))
        async with session.get(url, headers=headers) as response:
            response.raise_for_status()
            html = await response.text()
            soup = BeautifulSoup(html, 'html.parser')
            product_title = soup.find("span", {"id": "productTitle"})
            if product_title:
                product_name = product_title.get_text().strip().split('|')[0].strip()
                return (asin, product_name)
            else:
                alt_title = soup.find("h1", class_="a-size-large a-spacing-none")
                return (asin, alt_title.get_text().strip() if alt_title else asin)
    except Exception as e:
        print(f"ASIN {asin}请求失败: {str(e)}")
        return (asin, asin)

async def scrape_product_names_async(asins):
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8',
        'Accept-Language': 'en-US,en;q=0.5',
        'Referer': 'https://www.amazon.com/',
        'Upgrade-Insecure-Requests': '1',
        'Connection': 'keep-alive'
    }
    async with aiohttp.ClientSession() as session:
        tasks = [fetch_asin(session, asin, headers) for asin in asins]
        results = await asyncio.gather(*tasks)
        return dict(results)

Streamlit中调用方式:

import streamlit as st
import asyncio

asins = ['B0XXXXXXX', 'B0YYYYYYY']
if st.button('爬取商品名称'):
    with st.spinner('正在爬取...'):
        product_mapping = asyncio.run(scrape_product_names_async(asins))
    st.write(product_mapping)

2. 缓存爬取结果

用Streamlit的缓存装饰器@st.cache_data缓存已爬取的ASIN数据,避免重复请求:

@st.cache_data(ttl=86400)  # 缓存有效期1天
def scrape_product_names_cached(asins):
    # 放入优化后的爬虫逻辑
    return product_mapping

三、商品图片爬取优化

复用上述反爬和异步思路,针对图片节点调整选择器:

async def fetch_product_image(session, asin, headers):
    url = f"https://www.amazon.com/dp/{asin}"
    try:
        await asyncio.sleep(random.uniform(0.5,2))
        async with session.get(url, headers=headers) as response:
            response.raise_for_status()
            html = await response.text()
            soup = BeautifulSoup(html, 'html.parser')
            # 优先主图选择器
            img_tag = soup.find("img", {"id": "landingImage"})
            if img_tag:
                return (asin, img_tag.get('src'))
            else:
                # 备选图片节点
                alt_img = soup.find("div", class_="imgTagWrapper").find("img")
                return (asin, alt_img.get('src') if alt_img else None)
    except Exception as e:
        print(f"ASIN {asin}图片爬取失败: {str(e)}")
        return (asin, None)

内容的提问来源于stack exchange,提问作者Zhi Toung Yap

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 12:04:58