Python网页爬取:如何一次性获取分页隐藏的全部页面URL
需求说明
我需要爬取站点https://collegedunia.com/usa/biology-colleges的所有分页URL,但该站点分页仅显示部分页码(比如当前在page-16时,只能获取page-1、14、15、16的URL)。我试过两段Python代码:第一段只能拿到部分URL,第二段用递归能获取全部但实现太复杂,想要更简洁直接的一次性获取所有分页URL的方法。
我使用的代码
第一段代码(仅能获取部分URL)
from bs4 import BeautifulSoup import requests url = "https://collegedunia.com/usa/biology-colleges/page-16" headers = {'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/39.0.2171.95 Safari/537.36'} page = requests.Session().get(url, headers = headers) pageParse = BeautifulSoup(page.text, "lxml") soup = pageParse.find("ul", {'class':'jsx-3212034721 list-unstyled d-flex justify-content-center paginate text-center'}) soupATags = soup.find_all("a") paginationNos = [item.text for item in soupATags] pageURL = [f'https://collegedunia.com{item.get("href")}'for item in soupATags]
第二段代码(递归获取全部但实现复杂)
set1 = {"https://collegedunia.com/usa/biology-colleges"} blacklistSet2 = set() def funcCountryList(set2, set1): try: if len(set1)!=len(blacklistSet2): for url in set1: if not url in blacklistSet2: page = requests.Session().get(url, headers = headers) pageParse = BeautifulSoup(page.text, "lxml") soup = pageParse.find("ul", {'class':'jsx-3212034721 list-unstyled d-flex justify-content-center paginate text-center'}) soupATags = soup.find_all("a") paginationNos = [item.text for item in soupATags] pageURL = [f'https://collegedunia.com{item.get("href")}'for item in soupATags] for item in pageURL: set1.add(item) blacklistSet2.add(url) funcCountryList(blacklistSet2, set1) else: continue else: finalList.append(set1) return finalList except: return sorted(list(set1)) getURLS = funcCountryList(blacklistSet2, set1) print(getURLS)
简洁解决方案
利用该站点分页URL的固定规律,只需要一次请求获取总页数,就能直接构造所有分页URL,无需递归或多次请求:
from bs4 import BeautifulSoup import requests headers = {'User-Agent': 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_10_1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/39.0.2171.95 Safari/537.36'} base_url = "https://collegedunia.com/usa/biology-colleges" # 请求首页获取总页数 response = requests.get(base_url, headers=headers) soup = BeautifulSoup(response.text, "lxml") # 定位分页栏并提取所有数字页码 pagination_ul = soup.find("ul", {'class':'jsx-3212034721 list-unstyled d-flex justify-content-center paginate text-center'}) if pagination_ul: # 筛选出纯数字的页码文本 page_numbers = [tag.text for tag in pagination_ul.find_all("a") if tag.text.strip().isdigit()] if page_numbers: total_pages = int(max(page_numbers)) # 构造所有分页URL all_pages = [f"{base_url}/page-{page}" for page in range(1, total_pages + 1)] # 把首页(不带page参数的)加入列表 all_pages.insert(0, base_url) # 去重并排序 all_pages = sorted(list(set(all_pages))) print("所有分页URL:") for url in all_pages: print(url) else: print("未解析到有效页码") else: print("未找到分页组件")
方案说明
- 该站点的分页URL遵循
https://collegedunia.com/usa/biology-colleges/page-{页码}的规律,只要拿到总页数就能批量生成 - 只需请求一次首页(或任意包含总页数的页面),解析出最大页码后直接构造所有URL,比递归方案减少了大量重复请求,代码更简洁易维护
内容的提问来源于stack exchange,提问作者Vijith Kumar V
相关产品推荐
相关产品推荐

