Python3.4/BeautifulSoup爬虫脚本:如何实现日托网站多页爬取?
Hey there! Let's tackle this pagination issue step by step—first, we need to figure out how the website handles page turns, then adjust your script to loop through all pages.
解决跨页爬取的具体步骤
第一步:分析网站的翻页机制
首先打开浏览器的开发者工具(按F12),切换到Network标签页,然后点击网站的「下一页」按钮,观察新出现的请求:
- 如果是GET请求:URL里会有类似
?page=2、?p=3这样的参数,页码就是这里的数值;也可能用offset=25(对应第二页,因为每页25条)来控制位置。 - 如果是POST请求:需要看Request Payload里的参数,比如可能有
page_number: 2、current_page: 3这类字段,这些就是控制页码的关键。
这一步是核心,你得先确认网站传递页码的方式和参数名——不同网站的命名差异很大,必须对应上才能正确请求后续页面。
第二步:修改脚本实现循环爬取
假设我们已经明确了页码参数,这里给你两种常见场景的代码示例,你可以对应调整:
场景1:GET请求翻页
如果网站用URL参数传递页码,比如基础URL是https://example.com/daycares?page=1,脚本可以这样写:
import requests from bs4 import BeautifulSoup import time # 基础URL,页码会动态替换 base_url = "https://example.com/daycares?page={}" # 可以先爬第一页提取真实总页数,这里先暂时用你说的150+ total_pages = 156 for page_num in range(1, total_pages + 1): try: # 构造当前页的完整URL current_url = base_url.format(page_num) # 模拟浏览器请求,避免被反爬 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36" } response = requests.get(current_url, headers=headers) response.raise_for_status() # 检查请求是否成功 # 解析页面内容 soup = BeautifulSoup(response.text, 'html.parser') # 这里替换成你原来的解析逻辑,比如提取日托机构的名称、地址等 daycare_items = soup.find_all('div', class_='daycare-card') for item in daycare_items: name = item.find('h3').text.strip() address = item.find('p', class_='address').text.strip() # 可以把数据写入文件或数据库,这里先打印示例 print(f"第{page_num}页 | {name}: {address}") # 加延迟,避免请求过于频繁被封禁 time.sleep(1.5) except Exception as e: print(f"爬取第{page_num}页失败: {str(e)}") # 失败后可以暂停几秒再继续 time.sleep(3) continue
场景2:POST请求翻页
如果网站用POST表单传递页码,比如请求的Form Data里有page: 2,脚本调整如下:
import requests from bs4 import BeautifulSoup import time base_url = "https://example.com/daycares" total_pages = 156 for page_num in range(1, total_pages + 1): try: # 构造POST请求的参数,注意要把开发者工具里看到的所有必要参数都带上 payload = { "page": page_num, "per_page": 25, # 每页25条,和网站展示一致 # 可能还有其他参数,比如搜索条件、分类ID等,都要复制过来 } headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36" } response = requests.post(base_url, data=payload, headers=headers) response.raise_for_status() soup = BeautifulSoup(response.text, 'html.parser') # 执行你的解析逻辑 daycare_items = soup.find_all('div', class_='daycare-card') for item in daycare_items: # 提取信息... pass time.sleep(1.5) except Exception as e: print(f"爬取第{page_num}页失败: {str(e)}") time.sleep(3) continue
第三步:优化建议
- 动态获取总页数:不要硬写156,可以先爬第一页,找到页面上的总页数元素(比如「共156页」的文字),用BeautifulSoup提取数值,让脚本更灵活。
- Python3.4兼容性:注意
requests和beautifulsoup4的版本要适配Python3.4,比如安装时指定:pip install requests==2.27.1 beautifulsoup4==4.9.3 - 反爬应对:如果遇到验证码或IP封禁,可以考虑使用代理IP,或者增加请求间隔时间。
内容的提问来源于stack exchange,提问作者coder101
相关产品推荐
相关产品推荐

