You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取:如何一次性获取分页隐藏的全部页面URL

需求说明

我需要爬取站点https://collegedunia.com/usa/biology-colleges的所有分页URL,但该站点分页仅显示部分页码(比如当前在page-16时,只能获取page-1、14、15、16的URL)。我试过两段Python代码:第一段只能拿到部分URL,第二段用递归能获取全部但实现太复杂,想要更简洁直接的一次性获取所有分页URL的方法。

我使用的代码

第一段代码(仅能获取部分URL)

from bs4 import BeautifulSoup
import requests

url = "https://collegedunia.com/usa/biology-colleges/page-16"

headers = {'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/39.0.2171.95 Safari/537.36'}

page = requests.Session().get(url, headers = headers)

pageParse = BeautifulSoup(page.text, "lxml")

soup = pageParse.find("ul", {'class':'jsx-3212034721 list-unstyled d-flex justify-content-center paginate text-center'})

soupATags = soup.find_all("a")

paginationNos = [item.text for item in soupATags]

pageURL = [f'https://collegedunia.com{item.get("href")}'for item in soupATags]

第二段代码(递归获取全部但实现复杂)

set1 = {"https://collegedunia.com/usa/biology-colleges"}

blacklistSet2 = set()

def funcCountryList(set2, set1):

    try:

        if len(set1)!=len(blacklistSet2):

            for url in set1:

                if not url in blacklistSet2:

                    page = requests.Session().get(url, headers = headers)

                    pageParse = BeautifulSoup(page.text, "lxml")

                    soup = pageParse.find("ul", {'class':'jsx-3212034721 list-unstyled d-flex justify-content-center paginate text-center'})

                    soupATags = soup.find_all("a")

                    paginationNos = [item.text for item in soupATags]

                    pageURL = [f'https://collegedunia.com{item.get("href")}'for item in soupATags]

                    for item in pageURL:

                        set1.add(item)

                    blacklistSet2.add(url)

                    funcCountryList(blacklistSet2, set1)

                else:

                    continue
        else:

            finalList.append(set1)

            return finalList

    except:

        return sorted(list(set1))

getURLS = funcCountryList(blacklistSet2, set1)

print(getURLS)

简洁解决方案

利用该站点分页URL的固定规律,只需要一次请求获取总页数,就能直接构造所有分页URL,无需递归或多次请求:

from bs4 import BeautifulSoup
import requests

headers = {'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/39.0.2171.95 Safari/537.36'}
base_url = "https://collegedunia.com/usa/biology-colleges"

# 请求首页获取总页数
response = requests.get(base_url, headers=headers)
soup = BeautifulSoup(response.text, "lxml")

# 定位分页栏并提取所有数字页码
pagination_ul = soup.find("ul", {'class':'jsx-3212034721 list-unstyled d-flex justify-content-center paginate text-center'})
if pagination_ul:
    # 筛选出纯数字的页码文本
    page_numbers = [tag.text for tag in pagination_ul.find_all("a") if tag.text.strip().isdigit()]
    if page_numbers:
        total_pages = int(max(page_numbers))
        # 构造所有分页URL
        all_pages = [f"{base_url}/page-{page}" for page in range(1, total_pages + 1)]
        # 把首页(不带page参数的)加入列表
        all_pages.insert(0, base_url)
        # 去重并排序
        all_pages = sorted(list(set(all_pages)))
        
        print("所有分页URL:")
        for url in all_pages:
            print(url)
    else:
        print("未解析到有效页码")
else:
    print("未找到分页组件")

方案说明

  1. 该站点的分页URL遵循https://collegedunia.com/usa/biology-colleges/page-{页码}的规律,只要拿到总页数就能批量生成
  2. 只需请求一次首页(或任意包含总页数的页面),解析出最大页码后直接构造所有URL,比递归方案减少了大量重复请求,代码更简洁易维护

内容的提问来源于stack exchange,提问作者Vijith Kumar V

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 21:07:04