如何用Python爬取网页多页数据?单页爬取成功但多页失败
解决多页网页爬取问题
现有Python爬虫代码仅能获取目标网站第一页内容,该网站共40多页数据,预期抓取1406条数据,实际仅返回第一页的37条。原代码如下:
from bs4 import BeautifulSoup import requests from csv import writer url = "https://www.vivareal.com.br/venda/sp/sao-bernardo-do-campo/condominio_residencial/" page = requests.get(url) print(page) soup = BeautifulSoup(page.content, 'html.parser') lists = soup.find_all('article', class_="property-card__container js-property-card") with open('test.csv', 'w', encoding='utf8', newline='') as f: thewriter = writer(f) header = ['Title', 'Location', 'Price', 'Area'] thewriter.writerow(header) for list in lists: title = list.find('span', class_="js-card-title").text.replace('\n', '') location = list.find('span', class_="property-card__address").text.replace('\n', '') price = list.find('div', class_="js-property-card__price-small").text.replace('\n', '') area = list.find('span', class_="js-property-card-detail-area").text.replace('\n', '') info = [title, location, price, area] thewriter.writerow(info)
修改后的多页爬取代码
from bs4 import BeautifulSoup import requests from csv import writer # 基础URL,分页参数通过pagina字段传递 base_url = "https://www.vivareal.com.br/venda/sp/sao-bernardo-do-campo/condominio_residencial/" # 模拟浏览器请求头,避免被反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } # 先写入CSV表头 with open('test.csv', 'w', encoding='utf8', newline='') as f: thewriter = writer(f) header = ['Title', 'Location', 'Price', 'Area'] thewriter.writerow(header) # 遍历页码,假设最多45页,可根据实际情况调整 for page_num in range(1, 46): # 构造当前页URL current_url = f"{base_url}?pagina={page_num}" try: # 发送请求 page = requests.get(current_url, headers=headers) page.raise_for_status() # 检查请求是否成功 print(f"爬取第{page_num}页,状态码:{page.status_code}") soup = BeautifulSoup(page.content, 'html.parser') lists = soup.find_all('article', class_="property-card__container js-property-card") # 如果当前页没有数据,说明已到最后一页,终止循环 if not lists: print(f"第{page_num}页无数据,终止爬取") break # 追加写入当前页数据到CSV with open('test.csv', 'a', encoding='utf8', newline='') as f: thewriter = writer(f) for item in lists: # 处理可能的空值,避免AttributeError title = item.find('span', class_="js-card-title").text.replace('\n', '') if item.find('span', class_="js-card-title") else '无标题' location = item.find('span', class_="property-card__address").text.replace('\n', '') if item.find('span', class_="property-card__address") else '无地址' price = item.find('div', class_="js-property-card__price-small").text.replace('\n', '') if item.find('div', class_="js-property-card__price-small") else '无价格' area = item.find('span', class_="js-property-card-detail-area").text.replace('\n', '') if item.find('span', class_="js-property-card-detail-area") else '无面积' info = [title, location, price, area] thewriter.writerow(info) except requests.exceptions.RequestException as e: print(f"第{page_num}页爬取失败:{e}") continue
关键改动说明
- 分页URL构造:通过分析网站分页规则,使用
?pagina=参数拼接不同页码的URL,遍历1到45页(可根据实际总页数调整)。 - 请求头添加:加入
User-Agent模拟浏览器请求,避免被网站反爬机制拦截。 - 文件写入模式:先以
w模式写入表头,后续每页用a追加模式写入数据,防止覆盖之前的内容。 - 异常处理:捕获请求异常,遇到错误时跳过当前页继续爬取其他页,提高程序稳定性。
- 空值处理:对每个字段添加存在性检查,避免因页面元素缺失导致程序报错中断。
- 终止条件:如果某页没有抓取到数据,判定为已到最后一页,提前终止循环。
内容的提问来源于stack exchange,提问作者Vtcs
相关产品推荐
相关产品推荐

