如何使用Python爬取网页上「上一页」按钮的链接?
爬取动态生成的「上一页」链接解决方案
- 问题本质:目标页面的「上一页」按钮链接是JavaScript动态渲染的,BeautifulSoup仅能解析页面初始静态HTML,无法获取JS运行后生成的href属性,所以返回的
<a>标签为空。 - 可行方案:
- 方案一:使用浏览器模拟工具(如Selenium)获取渲染后的页面内容
以下是Python代码示例:from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options # 配置无头Chrome模式,避免弹出浏览器窗口 chrome_options = Options() chrome_options.add_argument("--headless=new") driver = webdriver.Chrome(options=chrome_options) # 目标页面URL target_url = "https://onlinehelp.prinect-lounge.com/Prinect_Color_Toolbox/Version2021/de_10/#t=Prinect%2Fmeasuring%2Fmeasuring-4.htm" driver.get(target_url) # 等待页面JS渲染完成(可根据实际情况调整等待时间) driver.implicitly_wait(3) # 定位「上一页」按钮并提取href back_btn = driver.find_element(By.ID, "browseSeqBack") prev_page_href = back_btn.get_attribute("href") print(prev_page_href) driver.quit() - 方案二:通过URL规则直接构造链接
观察页面URL规律:当前页面为measuring-4.htm,上一页是measuring-3.htm,可提取URL中文件名的数字部分,减1后拼接成新链接,无需爬取页面。比如用正则匹配数字:import re current_url = "https://onlinehelp.prinect-lounge.com/Prinect_Color_Toolbox/Version2021/de_10/#t=Prinect%2Fmeasuring%2Fmeasuring-4.htm" # 提取文件名中的数字 num_match = re.search(r"measuring-(\d+)\.htm", current_url) if num_match: current_num = int(num_match.group(1)) prev_num = current_num - 1 # 构造上一页链接 prev_href = current_url.replace(f"measuring-{current_num}.htm", f"measuring-{prev_num}.htm") # 处理锚点部分,转换为实际可访问的URL格式 prev_href = prev_href.replace("#t=", "") print(prev_href)
- 方案一:使用浏览器模拟工具(如Selenium)获取渲染后的页面内容
内容的提问来源于stack exchange,提问作者Jefferson
相关产品推荐
相关产品推荐

