如何提升Streamlit程序运行速度及优化亚马逊商品爬虫效果?
亚马逊商品名称爬虫优化方案
一、解决「获取数量少」的问题
1. 完善请求头,模拟真实浏览器
仅用User-Agent容易被亚马逊反爬机制识别,补充更多浏览器请求头字段,降低拦截概率:
headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.amazon.com/', 'Upgrade-Insecure-Requests': '1', 'Connection': 'keep-alive' }
2. 增加反爬应对机制
- 随机延迟:每次请求后随机等待1-3秒,避免请求频率触发限制
- 重试机制:请求失败或未获取标题时,重试1-2次,重试时可更换请求头或代理
- 代理IP轮换:针对大量ASIN爬取,用代理池切换IP,避免单IP被封禁
示例整合代码:
import time import random from requests.adapters import HTTPAdapter from urllib3.util.retry import Retry def create_session(): session = requests.Session() # 配置重试策略 retry = Retry(total=2, backoff_factor=1, status_forcelist=[429, 500, 502, 503, 504]) adapter = HTTPAdapter(max_retries=retry) session.mount('https://', adapter) return session def scrape_product_names(asins, csv_path='product_names.csv'): product_mapping = {} session = create_session() headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.amazon.com/', 'Upgrade-Insecure-Requests': '1', 'Connection': 'keep-alive' } for asin in asins: try: url = f"https://www.amazon.com/dp/{asin}" time.sleep(random.uniform(1,3)) # 随机延迟 response = session.get(url, headers=headers) response.raise_for_status() # 抛出HTTP错误 soup = BeautifulSoup(response.content, 'html.parser') # 优先尝试原选择器 product_title = soup.find("span", {"id": "productTitle"}) if product_title: product_name = product_title.get_text().strip().split('|')[0].strip() product_mapping[asin] = product_name else: # 备选选择器适配不同页面结构 alt_title = soup.find("h1", class_="a-size-large a-spacing-none") product_mapping[asin] = alt_title.get_text().strip() if alt_title else asin except Exception as e: print(f"处理ASIN {asin}出错: {str(e)}") product_mapping[asin] = asin return product_mapping
3. 适配多页面结构
部分商品页面(如第三方卖家页)的标题可能不在productTitle节点,增加备选选择器,比如span.a-text-title、h1.a-size-medium等,覆盖更多页面场景。
二、解决「运行耗时久」的问题
1. 改用异步请求
用aiohttp替代requests实现并发请求,大幅缩短大量ASIN的爬取总耗时:
import aiohttp import asyncio from bs4 import BeautifulSoup async def fetch_asin(session, asin, headers): url = f"https://www.amazon.com/dp/{asin}" try: await asyncio.sleep(random.uniform(0.5,2)) async with session.get(url, headers=headers) as response: response.raise_for_status() html = await response.text() soup = BeautifulSoup(html, 'html.parser') product_title = soup.find("span", {"id": "productTitle"}) if product_title: product_name = product_title.get_text().strip().split('|')[0].strip() return (asin, product_name) else: alt_title = soup.find("h1", class_="a-size-large a-spacing-none") return (asin, alt_title.get_text().strip() if alt_title else asin) except Exception as e: print(f"ASIN {asin}请求失败: {str(e)}") return (asin, asin) async def scrape_product_names_async(asins): headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': 'https://www.amazon.com/', 'Upgrade-Insecure-Requests': '1', 'Connection': 'keep-alive' } async with aiohttp.ClientSession() as session: tasks = [fetch_asin(session, asin, headers) for asin in asins] results = await asyncio.gather(*tasks) return dict(results)
Streamlit中调用方式:
import streamlit as st import asyncio asins = ['B0XXXXXXX', 'B0YYYYYYY'] if st.button('爬取商品名称'): with st.spinner('正在爬取...'): product_mapping = asyncio.run(scrape_product_names_async(asins)) st.write(product_mapping)
2. 缓存爬取结果
用Streamlit的缓存装饰器@st.cache_data缓存已爬取的ASIN数据,避免重复请求:
@st.cache_data(ttl=86400) # 缓存有效期1天 def scrape_product_names_cached(asins): # 放入优化后的爬虫逻辑 return product_mapping
三、商品图片爬取优化
复用上述反爬和异步思路,针对图片节点调整选择器:
async def fetch_product_image(session, asin, headers): url = f"https://www.amazon.com/dp/{asin}" try: await asyncio.sleep(random.uniform(0.5,2)) async with session.get(url, headers=headers) as response: response.raise_for_status() html = await response.text() soup = BeautifulSoup(html, 'html.parser') # 优先主图选择器 img_tag = soup.find("img", {"id": "landingImage"}) if img_tag: return (asin, img_tag.get('src')) else: # 备选图片节点 alt_img = soup.find("div", class_="imgTagWrapper").find("img") return (asin, alt_img.get('src') if alt_img else None) except Exception as e: print(f"ASIN {asin}图片爬取失败: {str(e)}") return (asin, None)
内容的提问来源于stack exchange,提问作者Zhi Toung Yap
相关产品推荐
相关产品推荐

