URL末尾带/profile致Yahoo Finance爬虫函数失效求助
问题原因及解决方法
核心原因分析
ua变量未定义:代码中请求头直接使用了ua但未提前赋值,导致请求参数错误,无法正常发起有效请求。- 反爬拦截机制:Yahoo Finance的/profile页面反爬策略更严格,仅简单设置User-Agent可能被识别为爬虫,返回验证页面、空白页等非目标内容。
- 动态内容渲染:部分页面数据通过JavaScript动态加载,
requests只能获取静态HTML源码,无法拿到渲染后的实际内容,导致正则匹配失效。
针对性解决步骤
1. 修复ua变量并完善请求头
先定义合法的User-Agent,同时补充更多请求头字段降低被拦截概率:
def get_name(ticker): import requests, re # 模拟Chrome浏览器的User-Agent ua = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' url = f'https://finance.yahoo.com/quote/{ticker}/profile' # 完善请求头,模拟正常浏览器访问逻辑 headers = { 'User-Agent': ua, 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8', 'Accept-Language': 'en-US,en;q=0.5', 'Referer': f'https://finance.yahoo.com/quote/{ticker}' } req = requests.get(url, headers=headers) # 先打印状态码确认请求是否成功 print(f"请求状态码: {req.status_code}") html = req.text try: name = re.search(r'(?<=<title>)(.*?)(?=\s\([A-Z]{3,4}\))', html) print(name.group() if name else "未匹配到公司名称") except Exception as e: print(f"匹配出错: {str(e)}")
2. 应对动态渲染问题
如果完善请求头后仍无法获取数据,说明页面依赖JavaScript渲染,需要用selenium模拟浏览器加载:
from selenium import webdriver from selenium.webdriver.chrome.options import Options import re def get_name(ticker): # 配置无头浏览器模式,避免弹出窗口 chrome_options = Options() chrome_options.add_argument('--headless=new') chrome_options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36') driver = webdriver.Chrome(options=chrome_options) url = f'https://finance.yahoo.com/quote/{ticker}/profile' try: driver.get(url) # 获取渲染后的完整页面源码 html = driver.page_source name = re.search(r'(?<=<title>)(.*?)(?=\s\([A-Z]{3,4}\))', html) print(name.group() if name else "未匹配到公司名称") except Exception as e: print(f"执行出错: {str(e)}") finally: driver.quit()
3. 调试建议
- 打印
req.text查看实际返回内容,如果是验证页面或空白页,说明被反爬拦截,可尝试添加代理或进一步调整请求头。 - 检查目标页面的title实际格式,若正则匹配失效,可根据真实HTML结构调整正则表达式。
内容的提问来源于stack exchange,提问作者Samantha Zweiban
相关产品推荐
相关产品推荐

