爬取网站无法获取JS渲染内容,已试selenium等工具,如何提取指定类名标签数据
解决方案
你现有代码的核心问题是逻辑不连贯,JS渲染完成后没有直接提取目标元素,后续新开的requests会话也未定义bibtex_id、xhr_url两个变量,自然无法获取到目标数据。可参考以下两种可行实现:
方案1:修正后的requests_html实现
import asyncio from requests_html import HTMLSession if asyncio.get_event_loop().is_running(): import nest_asyncio nest_asyncio.apply() session = HTMLSession() url = 'https://www.canlii.org/en/on/onltb/nav/date/2021/' r = session.get(url) # 等待2秒确保JS完全渲染,可根据网络情况调整时长 r.html.render(sleep=2) # 匹配所有class属性符合要求的标签 target_elements = r.html.find('.row.row-stripped.py-1.ml-0') result = [] for ele in target_elements: # 提取标签内部文本,可根据需求改为提取子标签属性等操作 result.append(ele.text) print(result)
方案2:更稳定的Selenium实现(适配反爬场景)
如果网站反爬规则拦截requests_html请求,可改用Selenium模拟真实浏览器行为获取数据:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.firefox.options import Options import time # 配置无头模式,运行时不弹出浏览器窗口 firefox_options = Options() firefox_options.add_argument('--headless') driver = webdriver.Firefox(options=firefox_options) # 隐式等待10秒,自动等待元素加载完成 driver.implicitly_wait(10) url = 'https://www.canlii.org/en/on/onltb/nav/date/2021/' driver.get(url) time.sleep(2) # 匹配所有目标class的元素 target_elements = driver.find_elements(By.CSS_SELECTOR, '.row.row-stripped.py-1.ml-0') result = [] for ele in target_elements: result.append(ele.text) print(result) driver.quit()
注意事项
- 运行前需安装对应依赖:requests_html方案执行
pip install requests-html nest_asyncio,Selenium方案执行pip install selenium并提前配置好匹配本地Firefox版本的geckodriver - 请勿高频发起请求,避免触发网站反爬机制导致IP被封禁
- 若目标元素未加载完全,可适当调高代码中sleep的等待时长
内容的提问来源于stack exchange,提问作者Ray R
相关产品推荐
相关产品推荐

