抓取a.href返回空值?如何获取商品名称与编码并遍历页面
问题解决方案
1. 核心语法错误修正
你写的 soup.find_all('a.href') 是语法错误:这行代码是在查找**标签名为a.href**的元素,而非带有href属性的<a>标签。正确写法应为:
# 查找所有带href属性的a标签 test = soup.find_all('a', href=True)
但即使修正该写法,你依然拿不到商品数据——该网站的商品列表是JavaScript动态渲染的,urlopen只能获取静态页面框架,无法加载动态生成的内容。
2. 改用Selenium获取动态渲染内容
你已导入Selenium,直接用它加载页面并获取渲染后的完整HTML:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time url = 'https://smws.com/all-whisky?sort=pricedesc&page=1&per-page=128' # 初始化Chrome浏览器(需确保chromedriver版本与浏览器匹配) driver = webdriver.Chrome() driver.get(url) # 等待商品列表加载完成 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, 'product-item')) ) # 留缓冲时间确保内容全部加载 time.sleep(2) # 获取渲染后的页面源码 soup = BeautifulSoup(driver.page_source, 'html.parser') driver.quit()
3. 提取商品名称与编码
观察页面结构,商品的编码和名称都包裹在.product-item元素内,可直接定位提取:
# 遍历所有商品项 products = soup.find_all('div', class_='product-item') for item in products: # 提取产品编码(如53.389) code = item.find('span', class_='product-code').text.strip() # 提取商品名称(如Goodnight and Joy) name = item.find('h3', class_='product-name').text.strip() print(f"编码: {code}, 名称: {name}")
4. 实现全页面遍历
通过修改URL中的page参数遍历所有页面,同时判断是否到达最后一页:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time base_url = 'https://smws.com/all-whisky?sort=pricedesc&page={}&per-page=128' current_page = 1 driver = webdriver.Chrome() while True: url = base_url.format(current_page) driver.get(url) # 等待页面加载,超时则判定为最后一页 try: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, 'product-item')) ) except: break time.sleep(2) soup = BeautifulSoup(driver.page_source, 'html.parser') products = soup.find_all('div', class_='product-item') if not products: break # 处理当前页商品数据 for item in products: code = item.find('span', class_='product-code').text.strip() name = item.find('h3', class_='product-name').text.strip() print(f"页码: {current_page}, 编码: {code}, 名称: {name}") current_page += 1 driver.quit()
注意事项
- 确保Chrome浏览器版本与chromedriver版本一致,否则Selenium无法正常运行。
- 可适当延长等待时间,避免因页面加载不完整导致数据抓取失败。
- 频繁请求可能触发网站反爬机制,建议添加更长的请求间隔或使用代理。
内容的提问来源于stack exchange,提问作者Ianucci
相关产品推荐
相关产品推荐

