Python网页爬取:同一页面存在两个下一页链接时如何仅获取一个
如何在Python爬取中仅获取页面中一个下一页链接(顶部/底部重复场景)
当页面顶部和底部都存在相同的下一页链接时,你可以通过以下几种简单方式只获取其中一个:
方法1:直接取第一个或最后一个匹配元素
页面中两个下一页链接的HTML结构通常一致(比如类名相同),可以先找到所有匹配的链接,再选择第一个或最后一个:
import requests from bs4 import BeautifulSoup # 目标页面URL target_url = "https://www.imdb.com/search/title/?groups=top_100&sort=user_rating,desc&ref_=adv_prv" # 模拟浏览器请求头,避免被反爬拦截 headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"} # 获取页面内容 response = requests.get(target_url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # 找到所有下一页链接 next_links = soup.find_all("a", class_="lister-page-next next-page") # 取第一个链接(对应顶部的下一页) if next_links: first_next = next_links[0]["href"] # 拼接成完整URL(IMDB返回的是相对路径) full_first_link = "https://www.imdb.com" + first_next print("顶部下一页链接:", full_first_link) # 取最后一个链接(对应底部的下一页) if next_links: last_next = next_links[-1]["href"] full_last_link = "https://www.imdb.com" + last_next print("底部下一页链接:", full_last_link)
方法2:通过父容器精准定位
如果需要明确指定取顶部还是底部的链接,可以利用它们所在的不同父容器来定位。以IMDB示例页面为例:
- 顶部下一页链接在
div.lister-top-right容器内 - 底部下一页链接在
div.desc容器内
用CSS选择器直接定位目标容器内的链接:
# 定位顶部的下一页链接 top_next_link = soup.select_one("div.lister-top-right a.lister-page-next.next-page") if top_next_link: full_top_link = "https://www.imdb.com" + top_next_link["href"] print("顶部下一页链接:", full_top_link) # 定位底部的下一页链接 bottom_next_link = soup.select_one("div.desc a.lister-page-next.next-page") if bottom_next_link: full_bottom_link = "https://www.imdb.com" + bottom_next_link["href"] print("底部下一页链接:", full_bottom_link)
注意事项
- 实际爬取时一定要携带合理的请求头,避免被网站反爬机制拦截
- 注意链接的拼接:如果返回的是相对路径,需要和网站域名组合成完整URL才能正常访问
内容的提问来源于stack exchange,提问作者Sven
相关产品推荐
相关产品推荐

