You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Aliexpress网页爬虫仅返回1条产品数据 如何获取全部产品信息

问题解决方法

你代码的核心问题是选择了错误的遍历对象:organic-list app-organic-search__list是整个产品列表的外层容器,单页面仅有1个该元素,所以for循环只会执行1次,仅能拿到第一条产品数据。

修改后的代码
from bs4 import BeautifulSoup
import requests

# 加请求头模拟浏览器访问,避免被网站反爬拦截返回异常页面
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
}
html_text = requests.get('https://dutch.alibaba.com/products/uhf_rfid_label.html?IndexArea=product_en&page=1', headers=headers).text
soup = BeautifulSoup(html_text, 'lxml')
# 先定位产品列表外层容器
list_container = soup.find('div', class_ ='organic-list app-organic-search__list')
# 从容器内找所有单个产品的节点,单个产品对应class为organic-gallery-offer__wrapper
producten = list_container.find_all('div', class_='organic-gallery-offer__wrapper')
for product in producten:
    # 增加判空逻辑,避免部分产品字段缺失导致代码报错
    product_naam = product.find('p', class_ = 'elements-title-normal__content large').text if product.find('p', class_ = 'elements-title-normal__content large') else "无产品名称"
    jaren_actief = product.find('span', class_ = 'seller-tag__year flex-no-shrink').text if product.find('span', class_ = 'seller-tag__year flex-no-shrink') else "无卖家经营年限信息"
    print(f"Product naam: {product_naam}")
    print(f"Jaren actief: {jaren_actief}\n")
补充说明
  • 如果后续网站更新调整了前端class名,可以打开浏览器F12开发者工具,在元素面板点选单个产品卡片,查看对应节点的class属性替换即可
  • 新增的判空逻辑可以避免部分特殊产品字段缺失导致的代码运行报错
  • 如果需要爬取多页数据,循环修改url中page参数的数值请求即可

内容的提问来源于stack exchange,提问作者Borealis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 18:21:03