You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python grequests和BeautifulSoup抓取商品页并提取UPC等详情

实现方案

我们可以在原有异步抓取逻辑的基础上,新增一轮单品页的异步请求,解析后通过单品页URL关联到原有列表信息即可,全程保持异步请求的高效性,完整可运行代码如下:

from bs4 import BeautifulSoup
import grequests
import pandas as pd

# STEP 1: 生成需要抓取的列表页URL
def get_urls():
    urls = []
    # 可根据需求修改抓取的页码范围
    for x in range(1,3):
        urls.append(f'https://books.toscrape.com/catalogue/page-{x}.html')
        print(f'生成列表页URL: 第{x}页')
    return urls

# STEP 2: 通用异步请求工具,列表页和单品页请求复用
def async_req(urls):
    reqs = [grequests.get(link) for link in urls]
    resp = grequests.map(reqs)
    # 自动过滤请求失败的响应
    valid_resp = [r for r in resp if r and r.status_code == 200]
    return valid_resp

# STEP 3: 解析列表页,获取基础图书信息和对应的单品页URL
def parse_list(resp):
    productlist = []
    for r in resp:
        sp = BeautifulSoup(r.text, 'lxml')
        items = sp.find_all('article', {'class': 'product_pod'})
        for item in items:
            product = {
                'title' : item.find('h3').text.strip(),
                'price': item.find('p', {'class': 'price_color'}).text.strip(),
                'single_url': 'https://books.toscrape.com/catalogue/' + item.find('a').attrs['href'],
                'thumbnail': 'https://books.toscrape.com/' + item.find('img', {'class': 'thumbnail'}).attrs['src'],
            }
            productlist.append(product)
            print(f'列表页解析完成,添加图书: {product["title"]}')
    return productlist

# STEP 4: 批量解析单品页,追加UPC等详情字段到对应图书条目
def parse_detail(productlist):
    # 提取所有单品页URL
    detail_urls = [p['single_url'] for p in productlist]
    # 异步请求所有单品页
    detail_resps = async_req(detail_urls)
    
    for resp in detail_resps:
        # 通过URL匹配到列表页解析出来的对应图书条目
        current_product = next(p for p in productlist if p['single_url'] == resp.url)
        sp = BeautifulSoup(resp.text, 'lxml')
        # 提取单品页的属性表
        info_rows = sp.select('table.table.table-striped tr')
        for row in info_rows:
            key = row.find('th').text.strip()
            value = row.find('td').text.strip()
            # 直接把属性追加到原条目,按需保留需要的字段即可,比如只保留UPC、库存
            current_product[key] = value
        # 额外提取商品描述,不需要可以删除
        desc_tag = sp.select_one('#product_description + p')
        if desc_tag:
            current_product['description'] = desc_tag.text.strip()
        print(f'单品页解析完成: {current_product["title"]}, UPC: {current_product["UPC"]}')
    
    return productlist

# 主执行逻辑
if __name__ == '__main__':
    list_urls = get_urls()
    list_resp = async_req(list_urls)
    product_list = parse_list(list_resp)
    # 新增单品页解析步骤
    full_product_list = parse_detail(product_list)
    # 导出CSV时会自动包含所有新增字段
    df = pd.DataFrame(full_product_list)
    df.to_csv('books_full.csv', index=False)
    print(f'全部数据导出完成,共{len(full_product_list)}条记录')

关键逻辑说明

  • 把原有异步请求逻辑抽成通用工具函数,列表页和详情页可以复用,同时加了请求有效性校验,避免无效响应导致脚本报错
  • 通过单品页URL作为关联键,自动把解析出来的UPC、库存、品类、评分等字段追加到原有的图书条目上,不需要手动做字段映射
  • 如果不需要详情页的全部字段,只需要在属性解析的逻辑里加筛选条件,只保留你需要的字段即可,导出CSV时会自动适配字段

内容的提问来源于stack exchange,提问作者MarkWP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 16:09:03