Python爬取GullAhmed网站无法获取全部分页数据仅能导出第一页求助
问题原因
- 代码仅硬编码了第1页的请求地址,没有编写分页遍历逻辑,仅发起了1次请求,因此只能采集到第1页数据
- 未携带标准浏览器请求头,网站反爬策略可能拦截后续分页的请求,即使加了循环也可能拿不到有效数据
- 缺少请求状态校验、异常兜底逻辑,后续分页请求失败或部分商品数据异常时,程序可能无提示终止
修复后的完整代码
import requests from bs4 import BeautifulSoup import pandas as pd import time # 全局配置 base_url = 'https://www.gulahmedshop.com/unstitched-fabric?p={}&product_list_limit=48' total_page = 25 suit = [] # 添加浏览器请求头,伪装成正常访问 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } for page in range(1, total_page + 1): # 构造当前分页地址,页码不足两位自动补0 current_url = base_url.format(f"{page:02d}") print(f"正在采集第{page}页:{current_url}") try: r = requests.get(current_url, headers=headers, timeout=10) # 校验请求是否成功 if r.status_code != 200: print(f"第{page}页请求失败,状态码:{r.status_code}") continue soup = BeautifulSoup(r.content, 'html.parser') content = soup.find_all('div', class_ = 'product-item-info') if not content: print(f"第{page}页未找到商品数据") continue for item in content: # 商品名称 try: name = item.find('a', class_ = 'product-item-link').text.strip() except: name = '' # 销售价 try: p_price = item.find('div',class_ ='price-box price-final_price') p_span_price = p_price.find('span',class_='price-container price-final_price tax weee') N_product_price = p_span_price.find('span', {"class" : 'price'}).text.strip() except AttributeError: N_product_price = '' # 原价 try: p_price = item.find('div',class_ ='price-box price-final_price') p_span_price = p_price.find('span',class_='old-price') old_product_price = p_span_price.find('span', {"class" : 'price'}).text.strip() except AttributeError: old_product_price = '' # 优惠标签 try: Offer= item.find('div', class_ = 'label-content').text.strip() except: Offer='' # 商品链接 try: links = item.find('a',{'class': 'product-item-link'})['href'] except: links = '' # 商品图片,加兜底避免未定义报错 images = '' image_list = item.find_all('img',{'class':'product-image-photo'},src=True) for i in image_list: if 'data:image' not in i['src']: images = i['src'] break fabric={ 'productname':name, 'Product_Sale_price': N_product_price, 'Product_Old_Price':old_product_price, 'Offer':Offer, 'product_image': images, 'links': links, } suit.append(fabric) # 每页采集完间隔2秒,避免触发反爬 time.sleep(2) except Exception as e: print(f"第{page}页采集出错,错误信息:{str(e)}") continue print(f"采集完成,共获取{len(suit)}条商品数据") df = pd.DataFrame(suit) print(df.head()) df.to_csv('E:/unstitched-fabric.csv', encoding='utf_8_sig') # 指定编码避免打开CSV乱码
补充说明
- 若运行后仍有分页请求被拦截,可适当延长
time.sleep的间隔时长 - 所有逻辑均添加了异常捕获,单页/单个商品数据异常不会中断整体采集流程
内容的提问来源于stack exchange,提问作者Muhammad Umer Lari
相关产品推荐
相关产品推荐

