BeautifulSoup耐克商品爬虫price变量始终取第一个值如何解决
异常成因
- 循环内获取价格时调用了全局
soup对象的find方法,soup对应整个页面的DOM树根节点,find方法默认返回全文档第一个匹配的元素,因此每次循环拿到的都是页面首个商品的价格。 - 商品名称、图片链接能正常拉取,是因为这两个字段的查找动作是在当前遍历到的单个商品子节点
pair下执行的,只会匹配当前商品范围内的元素。
修复步骤
- 首先补上遗漏的
csv模块导入语句,否则运行时会触发模块未找到报错 - 将价格字段的查找主体从全局
soup替换为当前循环的单个商品节点pair - 建议给csv文件指定utf-8编码,避免写入特殊字符时出现乱码
修正后完整代码如下:
from bs4 import BeautifulSoup import requests import csv # 补全缺失的依赖导入 source = requests.get('https://www.nike.com/fr/w/hommes-chaussures-nik1zy7ok').text soup = BeautifulSoup(source, 'lxml') csv_file = open('nikeshoes.csv', 'w', newline='', encoding='utf-8') csv_writer = csv.writer(csv_file) csv_writer.writerow(['name', 'price', 'image_scr']) for pair in soup.find_all('div', class_='product-card__body'): name = pair.a.text print(name) # 查找主体从soup改为pair,仅在当前商品DOM范围内匹配价格 price = pair.find('div', class_='product-price css-11s12ax is--current-price').text print(price) try: image_scr = pair.select_one('img.css-1fxh5tw.product-card__hero-image')['src'] except Exception as e: image_scr = None print(image_scr) print() csv_writer.writerow([name, price, image_scr]) csv_file.close()
内容的提问来源于stack exchange,提问作者William Gode
相关产品推荐
相关产品推荐

