如何用BeautifulSoup和Selenium爬取动态分页的ESG评级数据?
动态分页ESG评级数据爬取解决方案
问题描述
我正在使用BeautifulSoup和Selenium从Sustainalytics的ESG评级页面提取公司名称与ESG评级。当前代码可以正常提取第一页数据,但点击下一页时URL不会变化,页面也没有可用的分页href链接,尝试多种方案都无法实现多页爬取,求解决思路。
现有可运行的第一页提取代码如下:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.action_chains import ActionChains from selenium.webdriver.common.keys import Keys from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time import csv import requests from bs4 import BeautifulSoup as bs #%%% chrome_options = Options() chrome_options.add_argument("--headless") # Run Chrome in headless mode # Specify the path to the ChromeDriver executable webdriver_service = Service('/usr/local/bin/chromedriver') # Choose Chrome browser driver = webdriver.Chrome(service=webdriver_service, options=chrome_options) # URL of the website url = "https://www.sustainalytics.com/esg-ratings" driver.get(url) time.sleep(5) htlm_content = driver.page_source # Close the browser driver.quit() #%% # Open the webpage soup = bs(htlm_content, "html.parser") print(soup.prettify()) esg_ratings = soup.find('div',class_='col-md-8') #print(esg_ratings) #%% company_names = [] ratings = [] company_rows = soup.find_all(class_='company-row') #%% for company_row in company_rows: company_name_elem = company_row.find(class_='primary-color') if company_name_elem is not None: company_name = company_name_elem.get_text() esg_risk_rating_elem = company_row.find(class_='col-2') if esg_risk_rating_elem is not None: esg_risk_rating = esg_risk_rating_elem.get_text() company_names.append(company_names) ratings.append(ratings) print(f"Company: {company_name}") print(f"Rating: {esg_risk_rating}")
解决思路
1. 定位分页按钮并模拟点击
这类动态加载的页面,分页操作不会修改URL,需要直接定位「下一页」按钮元素,用Selenium模拟点击。可通过By.CSS_SELECTOR或By.XPATH定位,比如查找带有「Next」文本的按钮,或对应class属性的按钮。
2. 调整浏览器生命周期
原代码获取第一页源码后直接关闭浏览器,导致无法进行后续分页操作。需在完成所有爬取后再关闭浏览器,循环内每次点击分页后等待页面加载完成,再提取当前页数据。
3. 修复数据存储错误
原代码中company_names.append(company_names)和ratings.append(ratings)是逻辑错误,应将提取到的company_name和esg_risk_rating添加到列表,否则会导致列表嵌套混乱。
4. 优化动态加载等待与反爬
- 替换固定
time.sleep()为WebDriverWait,等待目标元素加载完成,提升效率与稳定性。 - 无头模式易被检测,建议添加窗口大小、User-Agent等参数,伪装正常浏览器请求。
修改后的完整代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time import csv from bs4 import BeautifulSoup as bs # 配置浏览器选项 chrome_options = Options() chrome_options.add_argument("--headless=new") # 新版无头模式更接近正常浏览器 chrome_options.add_argument("--window-size=1920,1080") chrome_options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") # 指定ChromeDriver路径 webdriver_service = Service('/usr/local/bin/chromedriver') driver = webdriver.Chrome(service=webdriver_service, options=chrome_options) url = "https://www.sustainalytics.com/esg-ratings" driver.get(url) # 初始化存储列表 company_names = [] ratings = [] try: # 循环爬取多页,这里设置爬取5页为例,可根据需求调整 for page in range(5): # 等待页面加载完成,等待公司行元素出现 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "company-row")) ) # 获取当前页源码 html_content = driver.page_source soup = bs(html_content, "html.parser") # 提取当前页数据 company_rows = soup.find_all(class_='company-row') for company_row in company_rows: company_name_elem = company_row.find(class_='primary-color') company_name = company_name_elem.get_text(strip=True) if company_name_elem else "N/A" esg_risk_rating_elem = company_row.find(class_='col-2') esg_risk_rating = esg_risk_rating_elem.get_text(strip=True) if esg_risk_rating_elem else "N/A" company_names.append(company_name) ratings.append(esg_risk_rating) print(f"Company: {company_name}") print(f"Rating: {esg_risk_rating}") # 尝试点击下一页按钮 try: # 定位下一页按钮,可根据实际页面元素调整定位方式 next_button = WebDriverWait(driver, 5).until( EC.element_to_be_clickable((By.CSS_SELECTOR, "button.pagination-next")) ) next_button.click() # 等待页面切换完成 time.sleep(2) except: print("已到达最后一页,停止爬取") break finally: # 完成爬取后关闭浏览器 driver.quit() # 可选:将数据保存到CSV with open('esg_ratings.csv', 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) writer.writerow(["Company Name", "ESG Risk Rating"]) for name, rating in zip(company_names, ratings): writer.writerow([name, rating])
代码说明
- 采用新版无头模式
--headless=new,降低被反爬检测的概率。 - 使用
WebDriverWait等待元素加载,替代固定延迟,提升爬取稳定性。 - 修复数据存储逻辑错误,正确保存提取的公司名称与评级。
- 添加异常处理,当无法点击下一页时自动终止爬取。
- 增加CSV保存功能,方便后续数据整理与分析。
内容的提问来源于stack exchange,提问作者nvtn
相关产品推荐
相关产品推荐

