You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬虫运行返回None,无法提取href链接问题求助

问题根因
  1. 目标页面的核心表格内容为JavaScript动态渲染生成,requests.get()仅能获取未执行JS的初始静态HTML,该内容中不存在你要查找的对应属性的元素,因此soup返回空列表,无输出。
  2. 代码逻辑顺序错误:Selenium Webdriver的初始化放在了BeautifulSoup页面处理之后,完全未用到Selenium的动态页面渲染能力,ChromeDriver配置无实际作用。
修正方案

首先确认已安装所需依赖:
pip install requests beautifulsoup4 selenium webdriver-manager

然后使用Selenium获取渲染完成的页面源码后再解析,参考代码如下:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
from bs4 import BeautifulSoup
import time

URL = 'https://www.rotowire.com/basketball/team.php?team=UTA'
# 初始化浏览器并加载页面
service = Service(ChromeDriverManager().install())
# 可选:添加无头模式,不弹出浏览器窗口
chrome_options = webdriver.ChromeOptions()
chrome_options.add_argument("--headless=new")

driver = webdriver.Chrome(service=service, options=chrome_options)
driver.get(URL)
# 等待动态内容加载完成
time.sleep(3)

# 读取渲染后的完整页面源码
soup = BeautifulSoup(driver.page_source, 'html.parser')

for item in soup.find_all(attrs={"aria-colindex": "3"}):
    href = item.get('href')
    # 过滤无href属性的元素
    if href:
        print(href)

# 关闭浏览器进程
driver.quit()
额外优化建议
  • 可将固定等待的time.sleep(3)替换为Selenium的显式等待,定位到目标元素加载完成后再执行解析,效率更高、兼容性更好
  • 可直接通过Selenium的元素定位方法直接提取href属性,无需额外引入BeautifulSoup做解析,减少依赖

内容的提问来源于stack exchange,提问作者Fraetos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 22:54:03