使用Python与Requests-HTML爬取JS渲染页面元素失败求助
解决Finviz页面爬取不到目标元素的问题
你的问题核心在于Finviz的反爬机制以及requests-html默认渲染配置不足,导致页面未完全加载或被识别为爬虫。以下是针对性的修复方案:
关键问题分析
- Finviz会检测请求的User-Agent,默认
requests-html的UA会被识别为爬虫,返回不完整页面 response.html.render()默认等待时间较短,可能JS还没完成渲染就获取了页面源码- 部分动态内容需要等待特定元素加载完成,而非单纯等待页面加载
修复后的代码
from bs4 import BeautifulSoup from requests_html import HTMLSession url = 'https://finviz.com/screener.ashx?v=111&f=sec_technology' session = HTMLSession() # 设置真实浏览器的User-Agent,避免被反爬识别 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } response = session.get(url, headers=headers) # 延长渲染等待时间,确保JS完全加载;同时禁用缓存强制刷新页面 response.html.render(timeout=20, wait=5, reload=True) # 方式1:直接用requests-html内置查找方法(无需BeautifulSoup) # tickers = response.html.find('.screener-link-primary') # 方式2:继续使用BeautifulSoup解析 soup = BeautifulSoup(response.html.html, "html.parser") tickers = soup.find_all('a', class_="screener-link-primary") # 提取并打印目标元素的文本内容 for ticker in tickers: print(ticker.get_text(strip=True))
额外注意事项
- 若仍无法获取内容,可尝试添加
scrolldown参数模拟页面滚动(如果目标内容是滚动加载的),例如response.html.render(scrolldown=3, sleep=2) - 频繁请求可能会被Finviz封禁IP,建议添加请求间隔,避免短时间内大量发起请求
- 若遇到人机验证,可能需要使用更复杂的工具(如Playwright)模拟真实浏览器交互
内容的提问来源于stack exchange,提问作者Jose Patiño
相关产品推荐
相关产品推荐

