You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改Python爬虫代码以抓取电商商品尺码、材质、图片等信息

twillmkt网站牛仔品类爬虫修改方案

原代码核心问题

  • 调用find_all('li').text直接取值报错:find_all()返回节点列表,无法直接读取text属性
  • 变量命名不规范:原Brand列表实际存储的是价格数据
  • 缺少新增字段的专属定位逻辑,没有拆分产品描述块里的多类信息

完整修改后代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

baseurl = 'https://twillmkt.com'
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.114 Safari/537.36'
}

# 抓取所有牛仔品类商品链接
r = requests.get('https://twillmkt.com/collections/denim', headers=headers)
soup = BeautifulSoup(r.content, 'html.parser')
tra = soup.find_all('div', class_='ProductItem__Wrapper')
productlinks = []
for links in tra:
    for link in links.find_all('a', href=True)[1:]:
        comp = baseurl + link['href']
        productlinks.append(comp)
# 链接去重
productlinks = list(set(productlinks))

# 初始化所有字段存储列表
Title = []
Price = []
Size = []
Product_Features = []
Material = []
Model_Size = []
Image = []

# 遍历单个商品页抓取字段
for link in productlinks:
    # 加异常处理避免单个页面失效导致整个程序终止
    try:
        r = requests.get(link, headers=headers)
        soup = BeautifulSoup(r.content, 'html.parser')
        # 标题
        title = soup.find('h1', class_='Product__Title').text.strip()
        Title.append(title)
        # 价格
        price = soup.find('span', class_='money').text.strip()
        Price.append(price)
        # 尺码
        size_list = [s.text.strip() for s in soup.find_all('label', class_='SizeSwatch') if s.text.strip()]
        Size.append(','.join(size_list))
        # 主图链接
        img_tag = soup.find('img', class_='Product__SlideshowImage')
        img_url = img_tag['src'] if img_tag else ''
        if img_url.startswith('//'):
            img_url = 'https:' + img_url
        Image.append(img_url)
        # 产品描述块拆分字段
        desc_items = soup.find('div', class_='Product__Description').find_all('li')
        desc_text = [item.text.strip() for item in desc_items]
        features = []
        material = ''
        model_info = ''
        for t in desc_text:
            t_low = t.lower()
            if 'material' in t_low or '材质' in t:
                material = t
            elif 'model' in t_low or '模特' in t or 'wear' in t_low:
                model_info = t
            else:
                features.append(t)
        Product_Features.append(';'.join(features))
        Material.append(material)
        Model_Size.append(model_info)
    except Exception as e:
        print(f"抓取{link}失败,错误信息:{e}")
        continue

# 构造数据表
df = pd.DataFrame({
    "商品标题": Title,
    "价格": Price,
    "尺码": Size,
    "产品特性": Product_Features,
    "材质": Material,
    "模特参数": Model_Size,
    "商品主图链接": Image
})
# 可直接导出为excel文件方便查看
df.to_excel('牛仔产品信息.xlsx', index=False)
print(df)

字段抓取逻辑说明

  • 尺码:直接提取页面上所有尺码选框的文本,多尺码用逗号拼接
  • 产品特性/材质/模特参数:统一从产品描述的列表项中提取,按关键词拆分对应字段
  • 商品图片:提取主图的CDN链接,自动补全http协议头
  • 新增了异常处理和链接去重逻辑,避免爬虫中途中断

内容的提问来源于stack exchange,提问作者Amen Aziz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 02:39:04