You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python 3.6同时提取多页面数据(无需URL列表)

多页数据提取的可行方案(无需完整URL列表)

嘿,完全没问题!针对你不想用完整URL列表的需求,这里有几种非常实用的Python 3.6多页数据提取方法,都是基于页面规律或自动抓取分页链接来实现的:

1. 基于分页参数自动生成URL

很多网站的分页逻辑是通过URL里的参数控制的(比如page=1、page=2),你只需要找到这个参数规律,就能循环生成所有分页的URL,不用手动列出来。

举个简单的实现示例:

import requests
from bs4 import BeautifulSoup
import time

# 替换成你要爬取的目标网站基础URL
base_url = "https://your-target-site.com/content?page="
# 可以先爬第一页获取最大页数,或者先设一个预估的最大值
max_pages = 15

for page_num in range(1, max_pages + 1):
    current_url = f"{base_url}{page_num}"
    try:
        # 模拟浏览器请求,避免被拦截
        headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"}
        response = requests.get(current_url, headers=headers)
        response.raise_for_status()  # 检查请求是否成功
        
        soup = BeautifulSoup(response.text, "html.parser")
        # 这里替换成你原本的单页数据提取逻辑
        target_items = soup.find_all("div", class_="content-item")
        for item in target_items:
            title = item.find("h2").text.strip()
            content = item.find("p").text.strip()
            print(f"第{page_num}页 - 标题:{title}")
            
        # 添加延时,避免请求过于频繁被封
        time.sleep(1)
    except Exception as e:
        print(f"爬取第{page_num}页失败:{str(e)}")
        continue

如果不确定最大页数,你可以先爬取第一页,解析分页栏里的最后一页数字,或者判断当返回的页面没有内容时自动停止循环。

2. 自动抓取「下一页」链接

有些网站的分页参数不明显,或者URL是动态生成的,这时候可以直接从页面中抓取「下一页」按钮的链接,循环爬取直到没有下一页为止。

示例代码:

import requests
from bs4 import BeautifulSoup
import time

current_url = "https://your-target-site.com/content"
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"}

while current_url:
    try:
        response = requests.get(current_url, headers=headers)
        response.raise_for_status()
        soup = BeautifulSoup(response.text, "html.parser")
        
        # 执行单页数据提取
        target_items = soup.find_all("div", class_="content-item")
        for item in target_items:
            title = item.find("h2").text.strip()
            print(f"标题:{title}")
        
        # 查找下一页链接(根据实际页面的按钮文本或class调整)
        next_btn = soup.find("a", text="下一页")  # 有些网站用class,比如class="next-page"
        if next_btn and "href" in next_btn.attrs:
            next_url = next_btn["href"]
            # 如果是相对路径,拼接成完整URL
            if not next_url.startswith("http"):
                current_url = f"https://your-target-site.com{next_url}"
            else:
                current_url = next_url
            time.sleep(1)
        else:
            # 没有下一页,终止循环
            current_url = None
            print("已爬取所有页面")
            
    except Exception as e:
        print(f"爬取失败:{str(e)}")
        current_url = None

3. 处理AJAX动态加载的页面

如果页面是通过滚动或点击触发AJAX请求加载数据(比如很多电商或社交网站),你需要抓包分析XHR请求的参数,直接请求数据接口。

示例(假设是GET接口):

import requests
import time

base_api_url = "https://your-target-site.com/api/get-content"
offset = 0
limit = 20  # 每页返回的数据条数,从抓包结果里查看

while True:
    params = {
        "offset": offset,
        "limit": limit,
        # 其他可能的参数,比如分类ID,从抓包结果里复制
        "category_id": 123
    }
    headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"}
    
    try:
        response = requests.get(base_api_url, params=params, headers=headers)
        response.raise_for_status()
        data = response.json()
        
        # 判断是否还有数据
        if not data.get("items"):
            print("没有更多数据了")
            break
        
        # 处理返回的JSON数据
        for item in data["items"]:
            print(f"标题:{item['title']}")
        
        offset += limit
        time.sleep(1)
    except Exception as e:
        print(f"请求接口失败:{str(e)}")
        break

一些注意事项

  • 一定要设置User-Agent请求头,模拟浏览器访问,避免被网站直接拦截。
  • 添加适当的延时(time.sleep()),不要短时间内发送大量请求,防止被封IP。
  • 处理异常情况(比如请求超时、页面解析失败),避免程序直接崩溃。
  • 遵守目标网站的robots.txt规则和使用条款,不要爬取敏感或禁止抓取的内容。

内容的提问来源于stack exchange,提问作者Razan Balatiah

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:57:29