网页爬取:无HTML页码时的分页处理(以Mecca护肤页为例)
解决Mecca护肤页面的分页爬取问题
问题说明
需要爬取Mecca护肤页面的产品品牌、名称及成分列表并保存到CSV,但原代码仅能爬取第一页,网站每页加载36个产品,URL以offset参数递增(如?offset=36、?offset=72),分页按钮无特定类属性,无法通过常规页码标识实现分页。
实现思路
利用URL的offset参数实现分页:
- 初始偏移量设为0,每次循环递增36
- 持续请求对应偏移量的页面,直到页面中不再包含产品元素时终止循环
- 同步完善CSV写入逻辑,将爬取到的结构化数据直接保存到文件
修改后的完整代码
import requests from bs4 import BeautifulSoup import csv baseurl = "https://www.mecca.com.au/skin-care/" offset = 0 # 初始化CSV文件 with open('mecca_skincare_products.csv', 'w', newline='', encoding='utf-8') as csvfile: fieldnames = ['品牌', '产品名称', '成分'] writer = csv.DictWriter(csvfile, fieldnames=fieldnames) writer.writeheader() while True: # 构造分页请求URL paginated_url = f"{baseurl}?offset={offset}" response = requests.get(paginated_url) soup = BeautifulSoup(response.content, "html.parser") products = soup.find_all('div', class_="grid-product-info") # 当前页面无产品时,终止循环 if not products: break # 收集当前页所有产品的完整链接 product_links = [] for item in products: for link in item.find_all('a', href=True): product_url = f"https://www.mecca.com.au{link['href']}" product_links.append(product_url) # 遍历产品链接,提取信息并写入CSV for link in product_links: product_res = requests.get(link) product_soup = BeautifulSoup(product_res.content, "html.parser") try: brand = product_soup.find('a', class_='css-1p371np-size5-size5-sansSerif-sansSerif-brandNameLink').text.strip() name = product_soup.find('span', class_='css-1noela6-paragraph-paragraph-sansSerif-sansSerif-productName').text.strip() # 提取成分内容 ingredients = "" for div in product_soup.find_all('div'): div_text = div.text.strip() if div_text.startswith('Ingredients:'): ingredients = div_text.split('Ingredients:', 1)[-1].strip() break # 写入CSV writer.writerow({'品牌': brand, '产品名称': name, '成分': ingredients}) print(f"已保存: {brand} - {name}") except AttributeError: # 处理页面元素缺失的异常情况 print(f"元素缺失,跳过链接: {link}") continue # 偏移量递增,进入下一页 offset += 36
关键优化点
- 修复原代码中链接拼接重复的问题(避免生成重复路径的无效链接)
- 新增完整的CSV写入逻辑,直接持久化爬取数据
- 添加异常处理,防止部分页面元素缺失导致程序崩溃
- 通过
offset参数循环遍历所有分页,实现全量产品爬取
内容的提问来源于stack exchange,提问作者user22276310
相关产品推荐
相关产品推荐

