You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬取速卖通无结果问题排查与解决求助

速卖通爬虫无法获取商品解决方案

问题概述

我是Web scraping新手,尝试编写Python代码爬取速卖通(AliExpress)网站,虽网传其对爬虫友好,但运行代码后始终显示“未找到商品”,无法获取商品URL。已添加超时、异常处理相关代码仍无效果,请求解决方案。

原代码如下:

import requests
from bs4 import BeautifulSoup
import csv
import time

base_url = "https://ar.aliexpress.com/?gatewayAdapt=glo2ara"
urls = [base_url]
products = []

while len(urls) != 0:
    current_url = urls.pop()

    try:
        response = requests.get(current_url, timeout=20)
        response.raise_for_status()
        soup = BeautifulSoup(response.content, "html.parser")

        link_elements = soup.select("a[href]")

        for link_element in link_elements:
            url = link_element['href']
            if base_url in url:
                urls.append(url)

                product = {}
                product["url"] = current_url

                # Corrected selector to target the image element
                image_element = soup.select_one(".product img")
                if image_element:
                    product["image"] = image_element["src"]
                else:
                    product["image"] = "Image not found"

                title_element = soup.select_one(".product_title")
                if title_element:
                    product["title"] = title_element.text 
                else:
                    product["title"] = "Title not found"

                # Add other fields as needed (e.g., price)

                products.append(product)

    except requests.exceptions.RequestException as e:
        print(f"An error occurred while processing {current_url}: {e}")

    time.sleep(0.1)
if products:
# Print the list of products
     print(products)

# Writing to CSV file
     csv_file_path = 'products.csv'
     with open(csv_file_path, 'w', newline='', encoding='utf-8') as csv_file:
         writer = csv.DictWriter(csv_file, fieldnames=products[0].keys())

    # Writing header
         writer.writeheader()

    # Writing rows
         writer.writerows(products)
 
     print(f"Data written to {csv_file_path}")
else:
    print("No products found.")

核心问题与解决步骤

  • 缺失请求头被反爬拦截:速卖通会校验请求中的User-Agent等标识,默认requests库的请求头会被识别为爬虫。需添加模拟浏览器的请求头,比如Chrome的UA。
  • 选择器与页面结构不匹配:原代码使用的.product、.product_title并非速卖通当前页面的商品元素类名,需要根据实际页面结构调整选择器。
  • URL筛选逻辑错误:原代码将所有包含base_url的链接都加入爬取队列,但大部分是导航、分类页面,需筛选包含item.htm的商品详情页URL。
  • 动态内容加载限制:速卖通部分商品数据由JavaScript动态渲染,静态requests请求可能无法获取完整内容,若基础调整后仍无效,可考虑使用Selenium/Playwright模拟浏览器渲染。

修正后的代码

import requests
from bs4 import BeautifulSoup
import csv
import time

# 模拟浏览器请求头
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Accept-Language": "ar-EG,ar;q=0.9,en-US;q=0.8,en;q=0.7"
}

base_url = "https://ar.aliexpress.com"
urls = [base_url]
products = []
visited_urls = set()  # 避免重复爬取同一页面

while urls:
    current_url = urls.pop()
    if current_url in visited_urls:
        continue
    visited_urls.add(current_url)

    try:
        response = requests.get(current_url, headers=headers, timeout=20)
        response.raise_for_status()
        soup = BeautifulSoup(response.content, "html.parser")

        # 提取所有链接,筛选商品详情页(包含item.htm)和站内其他页面
        for link in soup.select("a[href]"):
            href = link.get("href")
            if not href:
                continue
            # 处理相对链接
            full_url = href if href.startswith("http") else f"{base_url}{href}"
            # 筛选商品页或站内分类/导航页
            if base_url in full_url:
                if "item.htm" in full_url:
                    # 处理商品详情页
                    product = {"url": full_url}
                    # 访问商品详情页获取信息
                    try:
                        prod_response = requests.get(full_url, headers=headers, timeout=20)
                        prod_soup = BeautifulSoup(prod_response.content, "html.parser")
                        # 提取商品标题(根据当前页面结构调整)
                        title_elem = prod_soup.select_one("h1.product-title-text")
                        product["title"] = title_elem.get_text(strip=True) if title_elem else "Title not found"
                        # 提取商品图片(根据当前页面结构调整)
                        img_elem = prod_soup.select_one("img.magnifier-image")
                        product["image"] = img_elem.get("src") if img_elem else "Image not found"
                        products.append(product)
                    except Exception as e:
                        print(f"Failed to fetch product {full_url}: {e}")
                else:
                    # 添加非商品站内页面到队列(避免重复)
                    if full_url not in visited_urls:
                        urls.append(full_url)

        time.sleep(1)  # 延长休眠时间,避免频繁请求被封

    except requests.exceptions.RequestException as e:
        print(f"Error processing {current_url}: {e}")

if products:
    print(products)
    # 写入CSV
    csv_file_path = 'products.csv'
    with open(csv_file_path, 'w', newline='', encoding='utf-8') as csv_file:
        writer = csv.DictWriter(csv_file, fieldnames=products[0].keys())
        writer.writeheader()
        writer.writerows(products)
    print(f"Data saved to {csv_file_path}")
else:
    print("No products found.")

额外说明

  • 速卖通的页面结构可能随时变化,若选择器失效,需重新通过浏览器开发者工具查看元素类名或标签结构。
  • 频繁请求可能触发反爬机制,建议进一步增加休眠时间,或使用代理IP分散请求来源。
  • 若动态渲染内容无法通过静态请求获取,可替换为Selenium模拟浏览器操作,确保页面完全加载后再提取数据。

内容的提问来源于stack exchange,提问作者nadamalki

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 15:35:07