You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup和Selenium爬取动态分页的ESG评级数据?

动态分页ESG评级数据爬取解决方案

问题描述

我正在使用BeautifulSoup和Selenium从Sustainalytics的ESG评级页面提取公司名称与ESG评级。当前代码可以正常提取第一页数据,但点击下一页时URL不会变化,页面也没有可用的分页href链接,尝试多种方案都无法实现多页爬取,求解决思路。

现有可运行的第一页提取代码如下:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.action_chains import ActionChains
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time
import csv
import requests
from bs4 import BeautifulSoup as bs

#%%%
chrome_options = Options()
chrome_options.add_argument("--headless")
  # Run Chrome in headless mode

# Specify the path to the ChromeDriver executable
webdriver_service = Service('/usr/local/bin/chromedriver')

# Choose Chrome browser
driver = webdriver.Chrome(service=webdriver_service, options=chrome_options)

# URL of the website
url = "https://www.sustainalytics.com/esg-ratings"

driver.get(url)
time.sleep(5)

htlm_content = driver.page_source
# Close the browser
driver.quit()

#%%
# Open the webpage
soup = bs(htlm_content, "html.parser")
print(soup.prettify())
esg_ratings = soup.find('div',class_='col-md-8')
#print(esg_ratings)
#%%
company_names = []
ratings = []
company_rows = soup.find_all(class_='company-row')
#%%
for company_row in company_rows:
    company_name_elem = company_row.find(class_='primary-color')
    if company_name_elem is not None:
        company_name = company_name_elem.get_text()
    esg_risk_rating_elem = company_row.find(class_='col-2')
    if esg_risk_rating_elem is not None:
        esg_risk_rating = esg_risk_rating_elem.get_text()
    company_names.append(company_names)
    ratings.append(ratings)
    print(f"Company: {company_name}")
    print(f"Rating: {esg_risk_rating}")

解决思路

1. 定位分页按钮并模拟点击

这类动态加载的页面,分页操作不会修改URL,需要直接定位「下一页」按钮元素,用Selenium模拟点击。可通过By.CSS_SELECTOR或By.XPATH定位,比如查找带有「Next」文本的按钮,或对应class属性的按钮。

2. 调整浏览器生命周期

原代码获取第一页源码后直接关闭浏览器,导致无法进行后续分页操作。需在完成所有爬取后再关闭浏览器,循环内每次点击分页后等待页面加载完成,再提取当前页数据。

3. 修复数据存储错误

原代码中company_names.append(company_names)和ratings.append(ratings)是逻辑错误,应将提取到的company_name和esg_risk_rating添加到列表,否则会导致列表嵌套混乱。

4. 优化动态加载等待与反爬

  • 替换固定time.sleep()为WebDriverWait,等待目标元素加载完成,提升效率与稳定性。
  • 无头模式易被检测,建议添加窗口大小、User-Agent等参数,伪装正常浏览器请求。

修改后的完整代码

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time
import csv
from bs4 import BeautifulSoup as bs

# 配置浏览器选项
chrome_options = Options()
chrome_options.add_argument("--headless=new")  # 新版无头模式更接近正常浏览器
chrome_options.add_argument("--window-size=1920,1080")
chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")

# 指定ChromeDriver路径
webdriver_service = Service('/usr/local/bin/chromedriver')
driver = webdriver.Chrome(service=webdriver_service, options=chrome_options)

url = "https://www.sustainalytics.com/esg-ratings"
driver.get(url)

# 初始化存储列表
company_names = []
ratings = []

try:
    # 循环爬取多页,这里设置爬取5页为例,可根据需求调整
    for page in range(5):
        # 等待页面加载完成,等待公司行元素出现
        WebDriverWait(driver, 10).until(
            EC.presence_of_element_located((By.CLASS_NAME, "company-row"))
        )
        # 获取当前页源码
        html_content = driver.page_source
        soup = bs(html_content, "html.parser")
        
        # 提取当前页数据
        company_rows = soup.find_all(class_='company-row')
        for company_row in company_rows:
            company_name_elem = company_row.find(class_='primary-color')
            company_name = company_name_elem.get_text(strip=True) if company_name_elem else "N/A"
            
            esg_risk_rating_elem = company_row.find(class_='col-2')
            esg_risk_rating = esg_risk_rating_elem.get_text(strip=True) if esg_risk_rating_elem else "N/A"
            
            company_names.append(company_name)
            ratings.append(esg_risk_rating)
            print(f"Company: {company_name}")
            print(f"Rating: {esg_risk_rating}")
        
        # 尝试点击下一页按钮
        try:
            # 定位下一页按钮,可根据实际页面元素调整定位方式
            next_button = WebDriverWait(driver, 5).until(
                EC.element_to_be_clickable((By.CSS_SELECTOR, "button.pagination-next"))
            )
            next_button.click()
            # 等待页面切换完成
            time.sleep(2)
        except:
            print("已到达最后一页,停止爬取")
            break
finally:
    # 完成爬取后关闭浏览器
    driver.quit()

# 可选:将数据保存到CSV
with open('esg_ratings.csv', 'w', newline='', encoding='utf-8') as f:
    writer = csv.writer(f)
    writer.writerow(["Company Name", "ESG Risk Rating"])
    for name, rating in zip(company_names, ratings):
        writer.writerow([name, rating])

代码说明

  • 采用新版无头模式--headless=new,降低被反爬检测的概率。
  • 使用WebDriverWait等待元素加载,替代固定延迟,提升爬取稳定性。
  • 修复数据存储逻辑错误,正确保存提取的公司名称与评级。
  • 添加异常处理,当无法点击下一页时自动终止爬取。
  • 增加CSV保存功能,方便后续数据整理与分析。

内容的提问来源于stack exchange,提问作者nvtn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 17:37:50