You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Beautiful Soup解析该网站失效,求助Fitch Ratings网页爬取解决方案

问题原因

你遇到的情况是因为这个页面的内容是JavaScript动态渲染的——初始请求返回的静态HTML里没有目标数据,页面加载完成后,浏览器会通过JavaScript调用后端API拉取数据并渲染到页面上。所以直接用requests获取静态HTML,再用BeautifulSoup解析自然拿不到内容。

解决方案

方法1:直接爬取后端API接口(推荐)

这是效率最高的方式,无需渲染JS:

  • 打开浏览器开发者工具(按F12),切换到「Network」标签,刷新目标页面;
  • 在「XHR」或「Fetch」分类下,找到返回搜索结果的API请求(通常URL包含search、results这类关键词,响应格式为JSON);
  • 复制该API的请求URL、请求头(重点保留User-Agent、Referer,必要时带上Cookie);
  • 用requests库直接请求这个API,解析返回的JSON数据即可提取DATE、TITLE等字段。

示例代码(假设找到的API接口为https://www.fitchratings.com/search/results,参数需根据实际请求调整):

import requests

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Referer": "https://www.fitchratings.com/search/?expanded=racs&filter.country=&filter.language=English&filter.region=&filter.sector=&filter.sector=Structured%20Finance%3A%20RMBS&filter.topic=&isIdentifier=true&page=1&viewType=data"
}

params = {
    "page": 1,
    "filter.sector": "Structured Finance: RMBS",
    "isIdentifier": "true",
    # 其他从开发者工具复制的必要参数
}

response = requests.get("https://www.fitchratings.com/search/results", headers=headers, params=params)
data = response.json()

# 提取目标字段
for item in data["results"]:
    date = item.get("date")
    title = item.get("title")
    type_ = item.get("type")
    sector = item.get("sector")
    country = item.get("country")
    analysts = item.get("analysts")
    print(date, title, type_, sector, country, analysts)

方法2:使用无头浏览器渲染JavaScript

如果找不到API接口,可以用Selenium或Playwright这类工具模拟浏览器加载页面,等JS渲染完成后再提取内容:
示例代码(Selenium):

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup
import time

# 配置无头浏览器
chrome_options = Options()
chrome_options.add_argument("--headless=new")
chrome_options.add_argument("--disable-gpu")
chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")

driver = webdriver.Chrome(options=chrome_options)
url = "https://www.fitchratings.com/search/?expanded=racs&filter.country=&filter.language=English&filter.region=&filter.sector=&filter.sector=Structured%20Finance%3A%20RMBS&filter.topic=&isIdentifier=true&page=1&viewType=data"

driver.get(url)
# 等待页面渲染完成(可根据实际情况调整等待时间,或用显式等待)
time.sleep(3)

page_source = driver.page_source
soup = BeautifulSoup(page_source, "html.parser")

# 根据渲染后的HTML结构定位元素(需自行确认实际选择器)
results = soup.find_all("div", class_="search-result-item")
for item in results:
    date = item.find("span", class_="result-date").text.strip()
    title = item.find("h3", class_="result-title").text.strip()
    type_ = item.find("span", class_="result-type").text.strip()
    sector = item.find("span", class_="result-sector").text.strip()
    country = item.find("span", class_="result-country").text.strip()
    analysts = item.find("span", class_="result-analysts").text.strip()
    print(date, title, type_, sector, country, analysts)

driver.quit()
网站爬取可行性

该网站并非完全不可爬:

  1. 先查看网站的robots.txt(访问https://www.fitchratings.com/robots.txt),确认搜索页是否被禁止爬取;
  2. 爬取时需遵守爬虫礼仪:设置合理的请求间隔(比如每请求一次等待1-2秒),使用真实的User-Agent,避免短时间内大量请求导致IP被封禁;
  3. 当前搜索页的公开数据无需登录即可通过API或无头浏览器获取。

内容的提问来源于stack exchange,提问作者Tales_of_SS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 23:22:56