You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup无法提取动态网页<a>标签返回空结果的问题

问题描述

目标网页存在多个<a>标签,但使用BeautifulSoup调用findAll('a')或带属性筛选的方式均返回空结果,无法提取关联文本为"BT-Plenarprotokoll 20/86, S. 10313C"的目标<a>标签片段。

原因

该页面属于动态渲染页面:初始requests.get()获取的静态HTML中不包含实际页面内容,所有链接、文本都是通过JavaScript动态加载生成的,BeautifulSoup无法解析未加载的动态内容。

解决方案

方法1:用Selenium模拟浏览器渲染

Selenium会启动真实浏览器,等待页面完全加载后获取完整DOM内容,再用BeautifulSoup解析。

示例代码:

from bs4 import BeautifulSoup
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

url = "https://dip.bundestag.de/aktivität/Dr--Holger-Becker-MdB-SPD/1628877"

# 初始化Chrome浏览器(需确保chromedriver版本与浏览器匹配)
driver = webdriver.Chrome()
driver.get(url)

# 等待目标文本对应的元素加载完成,超时时间10秒
wait = WebDriverWait(driver, 10)
wait.until(EC.presence_of_element_located((By.XPATH, "//span[text()='BT-Plenarprotokoll 20/86, S. 10313C']")))

# 获取渲染后的完整页面源码
page_source = driver.page_source
driver.quit()

# 解析并提取目标a标签
soup = BeautifulSoup(page_source, 'html.parser')
target_span = soup.find("span", text="BT-Plenarprotokoll 20/86, S. 10313C")
if target_span:
    target_a = target_span.parent
    print("目标链接:", target_a['href'])

方法2:抓包获取API接口

打开浏览器开发者工具(F12),切换到「Network」标签页,刷新页面后筛选XHR/Fetch请求,找到返回页面数据的API接口,直接请求该接口获取结构化JSON数据,从中提取目标链接。这种方式无需解析HTML,效率更高。

额外提示
  • 静态解析工具(BeautifulSoup+requests)仅适用于内容直接写入初始HTML的页面,动态渲染页面必须依赖浏览器模拟或API请求。
  • 使用Selenium时可添加无头模式参数(options.add_argument('--headless=new')),避免弹出浏览器窗口。

内容的提问来源于stack exchange,提问作者jvqp

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 23:02:03