You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

爬取网站无法获取JS渲染内容,已试selenium等工具,如何提取指定类名标签数据

解决方案

你现有代码的核心问题是逻辑不连贯,JS渲染完成后没有直接提取目标元素,后续新开的requests会话也未定义bibtex_id、xhr_url两个变量,自然无法获取到目标数据。可参考以下两种可行实现:

方案1:修正后的requests_html实现

import asyncio
from requests_html import HTMLSession
if asyncio.get_event_loop().is_running():
    import nest_asyncio
    nest_asyncio.apply()

session = HTMLSession()
url = 'https://www.canlii.org/en/on/onltb/nav/date/2021/'
r = session.get(url)
# 等待2秒确保JS完全渲染,可根据网络情况调整时长
r.html.render(sleep=2)
# 匹配所有class属性符合要求的标签
target_elements = r.html.find('.row.row-stripped.py-1.ml-0')
result = []
for ele in target_elements:
    # 提取标签内部文本,可根据需求改为提取子标签属性等操作
    result.append(ele.text)

print(result)

方案2:更稳定的Selenium实现(适配反爬场景)

如果网站反爬规则拦截requests_html请求,可改用Selenium模拟真实浏览器行为获取数据:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.firefox.options import Options
import time

# 配置无头模式,运行时不弹出浏览器窗口
firefox_options = Options()
firefox_options.add_argument('--headless')
driver = webdriver.Firefox(options=firefox_options)
# 隐式等待10秒,自动等待元素加载完成
driver.implicitly_wait(10)

url = 'https://www.canlii.org/en/on/onltb/nav/date/2021/'
driver.get(url)
time.sleep(2)
# 匹配所有目标class的元素
target_elements = driver.find_elements(By.CSS_SELECTOR, '.row.row-stripped.py-1.ml-0')
result = []
for ele in target_elements:
    result.append(ele.text)

print(result)
driver.quit()

注意事项

  • 运行前需安装对应依赖:requests_html方案执行pip install requests-html nest_asyncio,Selenium方案执行pip install selenium并提前配置好匹配本地Firefox版本的geckodriver
  • 请勿高频发起请求,避免触发网站反爬机制导致IP被封禁
  • 若目标元素未加载完全,可适当调高代码中sleep的等待时长

内容的提问来源于stack exchange,提问作者Ray R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 10:45:07