You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python爬取网页多页数据?单页爬取成功但多页失败

解决多页网页爬取问题

现有Python爬虫代码仅能获取目标网站第一页内容,该网站共40多页数据,预期抓取1406条数据,实际仅返回第一页的37条。原代码如下:

from bs4 import BeautifulSoup
import requests
from csv import writer


url = "https://www.vivareal.com.br/venda/sp/sao-bernardo-do-campo/condominio_residencial/"
page = requests.get(url)
print(page)

soup = BeautifulSoup(page.content, 'html.parser')
lists = soup.find_all('article', class_="property-card__container js-property-card")

with open('test.csv', 'w', encoding='utf8', newline='') as f:
    thewriter = writer(f)
    header = ['Title', 'Location', 'Price', 'Area']
    thewriter.writerow(header)

    for list in lists:
        title = list.find('span', class_="js-card-title").text.replace('\n', '')
        location = list.find('span', class_="property-card__address").text.replace('\n', '')
        price = list.find('div', class_="js-property-card__price-small").text.replace('\n', '')
        area = list.find('span', class_="js-property-card-detail-area").text.replace('\n', '')

        info = [title, location, price, area]
        thewriter.writerow(info)

修改后的多页爬取代码

from bs4 import BeautifulSoup
import requests
from csv import writer

# 基础URL,分页参数通过pagina字段传递
base_url = "https://www.vivareal.com.br/venda/sp/sao-bernardo-do-campo/condominio_residencial/"
# 模拟浏览器请求头,避免被反爬拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36'
}

# 先写入CSV表头
with open('test.csv', 'w', encoding='utf8', newline='') as f:
    thewriter = writer(f)
    header = ['Title', 'Location', 'Price', 'Area']
    thewriter.writerow(header)

# 遍历页码,假设最多45页,可根据实际情况调整
for page_num in range(1, 46):
    # 构造当前页URL
    current_url = f"{base_url}?pagina={page_num}"
    try:
        # 发送请求
        page = requests.get(current_url, headers=headers)
        page.raise_for_status()  # 检查请求是否成功
        print(f"爬取第{page_num}页,状态码:{page.status_code}")

        soup = BeautifulSoup(page.content, 'html.parser')
        lists = soup.find_all('article', class_="property-card__container js-property-card")

        # 如果当前页没有数据,说明已到最后一页,终止循环
        if not lists:
            print(f"第{page_num}页无数据,终止爬取")
            break

        # 追加写入当前页数据到CSV
        with open('test.csv', 'a', encoding='utf8', newline='') as f:
            thewriter = writer(f)
            for item in lists:
                # 处理可能的空值,避免AttributeError
                title = item.find('span', class_="js-card-title").text.replace('\n', '') if item.find('span', class_="js-card-title") else '无标题'
                location = item.find('span', class_="property-card__address").text.replace('\n', '') if item.find('span', class_="property-card__address") else '无地址'
                price = item.find('div', class_="js-property-card__price-small").text.replace('\n', '') if item.find('div', class_="js-property-card__price-small") else '无价格'
                area = item.find('span', class_="js-property-card-detail-area").text.replace('\n', '') if item.find('span', class_="js-property-card-detail-area") else '无面积'

                info = [title, location, price, area]
                thewriter.writerow(info)
    except requests.exceptions.RequestException as e:
        print(f"第{page_num}页爬取失败:{e}")
        continue

关键改动说明

  • 分页URL构造:通过分析网站分页规则,使用?pagina=参数拼接不同页码的URL,遍历1到45页(可根据实际总页数调整)。
  • 请求头添加:加入User-Agent模拟浏览器请求,避免被网站反爬机制拦截。
  • 文件写入模式:先以w模式写入表头,后续每页用a追加模式写入数据,防止覆盖之前的内容。
  • 异常处理:捕获请求异常,遇到错误时跳过当前页继续爬取其他页,提高程序稳定性。
  • 空值处理:对每个字段添加存在性检查,避免因页面元素缺失导致程序报错中断。
  • 终止条件:如果某页没有抓取到数据,判定为已到最后一页,提前终止循环。

内容的提问来源于stack exchange,提问作者Vtcs

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 08:10:36