BeautifulSoup4解析器全部失效?速卖通爬取返回空列表求解决
解决速卖通爬取返回空列表问题
问题根源
- 动态内容渲染:速卖通商品数据通过JavaScript动态加载,
requests直接请求只能拿到未渲染的静态HTML,不含实际商品元素。 - 动态类名:你代码里硬编码的类名(如
manhattan--container--1lP57Ag)是平台动态生成的,每次请求都会变化,必然匹配失败。 - 反爬拦截:未伪装请求头的
requests请求易被识别为爬虫,返回的页面可能无有效内容。
解决方案
用Selenium模拟浏览器加载页面,获取完全渲染后的HTML,同时改用基于元素结构/属性的稳定选择器,避免依赖动态类名。
修改后的代码示例
import random from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup products = { 1: "Hoodies", 2: "Sunglasses", 3: "Couple-T-shirts", 4: "Wall-Stickers", 5: "Rugs", 6: "Dog-Bed", 7: "Claw-Cutter", 8: "Fur-Remover", 9: "Led-Keyboard", 10: "Wireless-Chargers", 11: "Powerbank", 12: "Game-Controller", 13: "Portable-Speakers", 14: "Scalp-Massager", 15: "Blackhead-Remover", 16: "Lash-Products", 17: "Makeup-Kit", 18: "Air-Tag-Tracker", 19: "Air-Purifiers", 20: "Pixelart", 21: "Yoga-Mats", 22: "Face-Masks", 23: "Fitness-Watches", 24: "Resistance-Bands", 25: "Air-Purifiers", 26: "Cell-Phone-Mounts", 27: "Wireless-Security-Cameras", 28: "Massage-Tools", 29: "Air-Purifiers", 30: "Eyeliner-Pencil", 31: "Water-Filters", 32: "Slow-Feeder-Dog-Bowls", 33: "Video-Doorbells", 34: "Solar-Outdoor-Lights", 35: "Phone-Grip", 36: "Slow-Feeder-Dog-Bowls", 37: "Pajamas", 38: "Skin-Care-Oil", 39: "Flasks", 40: "Monitor-Holders", 41: "Watches", 42: "Rings", 43: "Monitor-Holders", 44: "Nail-Polish", 45: "Rice-Cooker", } i = random.randint(1, 45) url = f'https://www.aliexpress.com/af/{products[i]}.html?spm=a2g0o.productlist.10000020&initiative_id=SB_20230106091400&dida=y&origin=n' # 配置Chrome无头模式,避免弹出浏览器窗口 chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-blink-features=AutomationControlled") chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=chrome_options) driver.get(url) # 等待页面加载完成(可根据实际情况调整等待时间) driver.implicitly_wait(10) # 获取渲染后的页面源码 page_source = driver.page_source soup = BeautifulSoup(page_source, 'lxml') # 使用更稳定的选择器,根据元素特征筛选商品卡片 product_cards = soup.find_all('a', attrs={'data-spm': True}) if product_cards: # 提取第一个商品链接 product_href = product_cards[0]['href'] productlinkstring = f'https://www.aliexpress.com{product_href}' # 提取商品图片 product_img = product_cards[0].find('img') productimagelinkstring = product_img['src'] if product_img else '无图片链接' # 提取价格 product_price = product_cards[0].find('div', class_=lambda x: x and 'price' in x) productpricestring = product_price.text.strip() if product_price else '无价格' # 提取评分 product_rating = product_cards[0].find('span', class_=lambda x: x and 'evaluation' in x) productratingsting = product_rating.text.strip() if product_rating else '无评分' print(productlinkstring) print(productimagelinkstring) print(productpricestring) print(productratingsting) else: print("未找到商品卡片") driver.quit()
注意事项
- 先安装依赖:
pip install selenium,确保Chrome浏览器版本与驱动版本匹配。 - 速卖通反爬机制会更新,选择器需根据页面结构变化调整。
- 添加随机延迟,避免频繁请求导致账号封禁。
内容的提问来源于stack exchange,提问作者OFENROHR100
相关产品推荐
相关产品推荐

