Beautiful Soup解析该网站失效,求助Fitch Ratings网页爬取解决方案
问题原因
你遇到的情况是因为这个页面的内容是JavaScript动态渲染的——初始请求返回的静态HTML里没有目标数据,页面加载完成后,浏览器会通过JavaScript调用后端API拉取数据并渲染到页面上。所以直接用requests获取静态HTML,再用BeautifulSoup解析自然拿不到内容。
解决方案
方法1:直接爬取后端API接口(推荐)
这是效率最高的方式,无需渲染JS:
- 打开浏览器开发者工具(按F12),切换到「Network」标签,刷新目标页面;
- 在「XHR」或「Fetch」分类下,找到返回搜索结果的API请求(通常URL包含
search、results这类关键词,响应格式为JSON); - 复制该API的请求URL、请求头(重点保留
User-Agent、Referer,必要时带上Cookie); - 用
requests库直接请求这个API,解析返回的JSON数据即可提取DATE、TITLE等字段。
示例代码(假设找到的API接口为https://www.fitchratings.com/search/results,参数需根据实际请求调整):
import requests headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Referer": "https://www.fitchratings.com/search/?expanded=racs&filter.country=&filter.language=English&filter.region=&filter.sector=&filter.sector=Structured%20Finance%3A%20RMBS&filter.topic=&isIdentifier=true&page=1&viewType=data" } params = { "page": 1, "filter.sector": "Structured Finance: RMBS", "isIdentifier": "true", # 其他从开发者工具复制的必要参数 } response = requests.get("https://www.fitchratings.com/search/results", headers=headers, params=params) data = response.json() # 提取目标字段 for item in data["results"]: date = item.get("date") title = item.get("title") type_ = item.get("type") sector = item.get("sector") country = item.get("country") analysts = item.get("analysts") print(date, title, type_, sector, country, analysts)
方法2:使用无头浏览器渲染JavaScript
如果找不到API接口,可以用Selenium或Playwright这类工具模拟浏览器加载页面,等JS渲染完成后再提取内容:
示例代码(Selenium):
from selenium import webdriver from selenium.webdriver.chrome.options import Options from bs4 import BeautifulSoup import time # 配置无头浏览器 chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--disable-gpu") chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=chrome_options) url = "https://www.fitchratings.com/search/?expanded=racs&filter.country=&filter.language=English&filter.region=&filter.sector=&filter.sector=Structured%20Finance%3A%20RMBS&filter.topic=&isIdentifier=true&page=1&viewType=data" driver.get(url) # 等待页面渲染完成(可根据实际情况调整等待时间,或用显式等待) time.sleep(3) page_source = driver.page_source soup = BeautifulSoup(page_source, "html.parser") # 根据渲染后的HTML结构定位元素(需自行确认实际选择器) results = soup.find_all("div", class_="search-result-item") for item in results: date = item.find("span", class_="result-date").text.strip() title = item.find("h3", class_="result-title").text.strip() type_ = item.find("span", class_="result-type").text.strip() sector = item.find("span", class_="result-sector").text.strip() country = item.find("span", class_="result-country").text.strip() analysts = item.find("span", class_="result-analysts").text.strip() print(date, title, type_, sector, country, analysts) driver.quit()
网站爬取可行性
该网站并非完全不可爬:
- 先查看网站的
robots.txt(访问https://www.fitchratings.com/robots.txt),确认搜索页是否被禁止爬取; - 爬取时需遵守爬虫礼仪:设置合理的请求间隔(比如每请求一次等待1-2秒),使用真实的
User-Agent,避免短时间内大量请求导致IP被封禁; - 当前搜索页的公开数据无需登录即可通过API或无头浏览器获取。
内容的提问来源于stack exchange,提问作者Tales_of_SS
相关产品推荐
相关产品推荐

