如何使用Python grequests和BeautifulSoup抓取商品页并提取UPC等详情
实现方案
我们可以在原有异步抓取逻辑的基础上,新增一轮单品页的异步请求,解析后通过单品页URL关联到原有列表信息即可,全程保持异步请求的高效性,完整可运行代码如下:
from bs4 import BeautifulSoup import grequests import pandas as pd # STEP 1: 生成需要抓取的列表页URL def get_urls(): urls = [] # 可根据需求修改抓取的页码范围 for x in range(1,3): urls.append(f'https://books.toscrape.com/catalogue/page-{x}.html') print(f'生成列表页URL: 第{x}页') return urls # STEP 2: 通用异步请求工具,列表页和单品页请求复用 def async_req(urls): reqs = [grequests.get(link) for link in urls] resp = grequests.map(reqs) # 自动过滤请求失败的响应 valid_resp = [r for r in resp if r and r.status_code == 200] return valid_resp # STEP 3: 解析列表页,获取基础图书信息和对应的单品页URL def parse_list(resp): productlist = [] for r in resp: sp = BeautifulSoup(r.text, 'lxml') items = sp.find_all('article', {'class': 'product_pod'}) for item in items: product = { 'title' : item.find('h3').text.strip(), 'price': item.find('p', {'class': 'price_color'}).text.strip(), 'single_url': 'https://books.toscrape.com/catalogue/' + item.find('a').attrs['href'], 'thumbnail': 'https://books.toscrape.com/' + item.find('img', {'class': 'thumbnail'}).attrs['src'], } productlist.append(product) print(f'列表页解析完成,添加图书: {product["title"]}') return productlist # STEP 4: 批量解析单品页,追加UPC等详情字段到对应图书条目 def parse_detail(productlist): # 提取所有单品页URL detail_urls = [p['single_url'] for p in productlist] # 异步请求所有单品页 detail_resps = async_req(detail_urls) for resp in detail_resps: # 通过URL匹配到列表页解析出来的对应图书条目 current_product = next(p for p in productlist if p['single_url'] == resp.url) sp = BeautifulSoup(resp.text, 'lxml') # 提取单品页的属性表 info_rows = sp.select('table.table.table-striped tr') for row in info_rows: key = row.find('th').text.strip() value = row.find('td').text.strip() # 直接把属性追加到原条目,按需保留需要的字段即可,比如只保留UPC、库存 current_product[key] = value # 额外提取商品描述,不需要可以删除 desc_tag = sp.select_one('#product_description + p') if desc_tag: current_product['description'] = desc_tag.text.strip() print(f'单品页解析完成: {current_product["title"]}, UPC: {current_product["UPC"]}') return productlist # 主执行逻辑 if __name__ == '__main__': list_urls = get_urls() list_resp = async_req(list_urls) product_list = parse_list(list_resp) # 新增单品页解析步骤 full_product_list = parse_detail(product_list) # 导出CSV时会自动包含所有新增字段 df = pd.DataFrame(full_product_list) df.to_csv('books_full.csv', index=False) print(f'全部数据导出完成,共{len(full_product_list)}条记录')
关键逻辑说明
- 把原有异步请求逻辑抽成通用工具函数,列表页和详情页可以复用,同时加了请求有效性校验,避免无效响应导致脚本报错
- 通过单品页URL作为关联键,自动把解析出来的UPC、库存、品类、评分等字段追加到原有的图书条目上,不需要手动做字段映射
- 如果不需要详情页的全部字段,只需要在属性解析的逻辑里加筛选条件,只保留你需要的字段即可,导出CSV时会自动适配字段
内容的提问来源于stack exchange,提问作者MarkWP
相关产品推荐
相关产品推荐

