You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python爬取网页上「上一页」按钮的链接?

爬取动态生成的「上一页」链接解决方案
  • 问题本质:目标页面的「上一页」按钮链接是JavaScript动态渲染的,BeautifulSoup仅能解析页面初始静态HTML,无法获取JS运行后生成的href属性,所以返回的<a>标签为空。
  • 可行方案:
    • 方案一:使用浏览器模拟工具(如Selenium)获取渲染后的页面内容
      以下是Python代码示例:
      from selenium import webdriver
      from selenium.webdriver.common.by import By
      from selenium.webdriver.chrome.options import Options
      
      # 配置无头Chrome模式,避免弹出浏览器窗口
      chrome_options = Options()
      chrome_options.add_argument("--headless=new")
      driver = webdriver.Chrome(options=chrome_options)
      
      # 目标页面URL
      target_url = "https://onlinehelp.prinect-lounge.com/Prinect_Color_Toolbox/Version2021/de_10/#t=Prinect%2Fmeasuring%2Fmeasuring-4.htm"
      driver.get(target_url)
      # 等待页面JS渲染完成(可根据实际情况调整等待时间)
      driver.implicitly_wait(3)
      
      # 定位「上一页」按钮并提取href
      back_btn = driver.find_element(By.ID, "browseSeqBack")
      prev_page_href = back_btn.get_attribute("href")
      print(prev_page_href)
      
      driver.quit()
      
    • 方案二:通过URL规则直接构造链接
      观察页面URL规律:当前页面为measuring-4.htm,上一页是measuring-3.htm,可提取URL中文件名的数字部分,减1后拼接成新链接,无需爬取页面。比如用正则匹配数字:
      import re
      
      current_url = "https://onlinehelp.prinect-lounge.com/Prinect_Color_Toolbox/Version2021/de_10/#t=Prinect%2Fmeasuring%2Fmeasuring-4.htm"
      # 提取文件名中的数字
      num_match = re.search(r"measuring-(\d+)\.htm", current_url)
      if num_match:
          current_num = int(num_match.group(1))
          prev_num = current_num - 1
          # 构造上一页链接
          prev_href = current_url.replace(f"measuring-{current_num}.htm", f"measuring-{prev_num}.htm")
          # 处理锚点部分,转换为实际可访问的URL格式
          prev_href = prev_href.replace("#t=", "")
          print(prev_href)
      

内容的提问来源于stack exchange,提问作者Jefferson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 02:54:53