You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取:同一页面存在两个下一页链接时如何仅获取一个

如何在Python爬取中仅获取页面中一个下一页链接(顶部/底部重复场景)

当页面顶部和底部都存在相同的下一页链接时,你可以通过以下几种简单方式只获取其中一个:

方法1:直接取第一个或最后一个匹配元素

页面中两个下一页链接的HTML结构通常一致(比如类名相同),可以先找到所有匹配的链接,再选择第一个或最后一个:

import requests
from bs4 import BeautifulSoup

# 目标页面URL
target_url = "https://www.imdb.com/search/title/?groups=top_100&sort=user_rating,desc&ref_=adv_prv"
# 模拟浏览器请求头,避免被反爬拦截
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}

# 获取页面内容
response = requests.get(target_url, headers=headers)
soup = BeautifulSoup(response.text, "html.parser")

# 找到所有下一页链接
next_links = soup.find_all("a", class_="lister-page-next next-page")

# 取第一个链接(对应顶部的下一页)
if next_links:
    first_next = next_links[0]["href"]
    # 拼接成完整URL(IMDB返回的是相对路径)
    full_first_link = "https://www.imdb.com" + first_next
    print("顶部下一页链接:", full_first_link)

# 取最后一个链接(对应底部的下一页)
if next_links:
    last_next = next_links[-1]["href"]
    full_last_link = "https://www.imdb.com" + last_next
    print("底部下一页链接:", full_last_link)

方法2:通过父容器精准定位

如果需要明确指定取顶部还是底部的链接,可以利用它们所在的不同父容器来定位。以IMDB示例页面为例:

  • 顶部下一页链接在div.lister-top-right容器内
  • 底部下一页链接在div.desc容器内

用CSS选择器直接定位目标容器内的链接:

# 定位顶部的下一页链接
top_next_link = soup.select_one("div.lister-top-right a.lister-page-next.next-page")
if top_next_link:
    full_top_link = "https://www.imdb.com" + top_next_link["href"]
    print("顶部下一页链接:", full_top_link)

# 定位底部的下一页链接
bottom_next_link = soup.select_one("div.desc a.lister-page-next.next-page")
if bottom_next_link:
    full_bottom_link = "https://www.imdb.com" + bottom_next_link["href"]
    print("底部下一页链接:", full_bottom_link)

注意事项

  • 实际爬取时一定要携带合理的请求头,避免被网站反爬机制拦截
  • 注意链接的拼接:如果返回的是相对路径,需要和网站域名组合成完整URL才能正常访问

内容的提问来源于stack exchange,提问作者Sven

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 14:25:10