基于Selenium和Python的Coursera分页爬取问题求助
Coursera个人成就页分页爬取修复方案(Selenium + Python)
问题概述
- 爬取自身Coursera成就页面(已确认robots.txt允许),单页内容可正常爬取,但分页切换后无法执行后续爬取,
get_pages()方法存在逻辑缺陷 - 现状:打印显示进入第2页,但无爬取动作;CAPTCHA需手动处理
- 需求:实现自动分页切换、多页内容爬取并写入文件
核心问题分析
- 原
get_pages()逻辑错误:仅通过XPATH查找分页元素,未执行点击跳转操作,且循环逻辑仅处理初始页,无法遍历多页 - 元素选择器未重置:
selector_counter在每页爬取后未重置,导致第二页开始查找的元素位置偏移 - 硬编码选择器不稳定:依赖
:nth-child的CSS选择器易受页面结构变化影响,容错性差 - 固定延迟不可靠:使用
time.sleep()等待页面加载,不如显式等待灵活可靠
修复后的完整代码
import time import secret from selenium import webdriver from selenium.webdriver.common.by import By from selenium.common.exceptions import NoSuchElementException, ElementClickInterceptedException from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC template = 'Certificate Name: {}\nCredential ID: {}\n' class CourseraScraper: def __init__ (self, url: str, username: str, password: str): self.url = url self.username = username self.password = password self.browser = webdriver.Firefox() self.wait = WebDriverWait(self.browser, 10) # 显式等待实例 def login(self): self.browser.get(self.url) print('访问登录页面...') # 显式等待用户名输入框加载 username_input = self.wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, '#email'))) username_input.send_keys(self.username) print('输入用户名...') pwd_input = self.wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, '#password'))) pwd_input.send_keys(self.password) print('输入密码...') login_button = self.wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, 'button[type="submit"]'))) login_button.click() print('点击登录按钮...') print('请手动完成CAPTCHA验证,完成后程序将继续...') # 等待用户手动完成CAPTCHA,直到页面跳转到主页 self.wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, '.cds-Avatar-initial'))) print('CAPTCHA验证完成,进入主页...') def click_accomplishments(self): print('打开用户下拉菜单...') dropdown = self.wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, '.cds-Avatar-initial'))) dropdown.click() print('点击"成就"链接...') # 替换更可靠的选择器,避免依赖:nth-child accomplishments_link = self.wait.until(EC.element_to_be_clickable((By.XPATH, '//a[contains(text(), "Accomplishments")]'))) accomplishments_link.click() # 等待成就页面加载完成 self.wait.until(EC.presence_of_element_located((By.CLASS_NAME, 'rc-AccomplishmentCard'))) print('进入成就页面...') def get_pages(self): current_page = 1 while True: print(f'开始爬取第 {current_page} 页...') self.scrape_page() # 尝试查找下一页按钮 try: # 定位下一页按钮(避免硬编码ID,用更通用的选择器) next_page_btn = self.wait.until(EC.element_to_be_clickable( (By.XPATH, '//button[contains(@aria-label, "Next page")]') )) # 检查按钮是否禁用(如果是最后一页,按钮会有disabled属性) if 'disabled' in next_page_btn.get_attribute('class'): print('已到达最后一页,停止爬取...') break next_page_btn.click() # 等待页面切换完成 self.wait.until(EC.staleness_of(self.browser.find_element(By.CLASS_NAME, 'rc-AccomplishmentCard'))) current_page += 1 time.sleep(1) # 短暂等待页面稳定 except (NoSuchElementException, ElementClickInterceptedException): print('未找到下一页按钮,爬取结束...') break print(f'爬取完成,共爬取 {current_page} 页') def scrape_page(self): # 直接获取所有成就卡片,避免依赖:nth-child accomplishment_cards = self.wait.until(EC.presence_of_all_elements_located( (By.CLASS_NAME, 'rc-AccomplishmentCard') )) print(f'当前页找到 {len(accomplishment_cards)} 个成就卡片') for card in accomplishment_cards: try: h3_text = card.find_element(By.TAG_NAME, 'h3').text link = card.find_element(By.CSS_SELECTOR, 'a[href*="/verify/"]') credential_id = link.get_attribute('href').rsplit('/', 1)[-1] formatted_text = template.format(h3_text, credential_id) print(formatted_text) self.write_text_to_file(formatted_text) except Exception as e: print(f'处理卡片时出错: {e}') continue def write_text_to_file(self, write_text: str): with open('coursera_accomplishments.txt', 'a+', encoding='utf-8') as f: f.write(write_text + '\n') def close(self): self.browser.quit() print('浏览器已关闭,任务完成!') if __name__ == '__main__': url = secret.url username = secret.username password = secret.password scraper = CourseraScraper(url, username, password) try: scraper.login() scraper.click_accomplishments() scraper.get_pages() finally: scraper.close()
关键修改说明
重构分页逻辑:
- 替换原
get_pages()的错误逻辑,改为通过"下一页"按钮判断是否还有后续页面,自动遍历所有分页 - 使用显式等待判断页面切换完成,避免固定延迟
- 替换原
优化元素选择:
- 移除依赖
:nth-child的不稳定选择器,改为直接获取所有成就卡片,遍历处理 - 用文本内容定位"成就"链接,避免依赖列表项索引
- 移除依赖
替换固定延迟为显式等待:
- 初始化
WebDriverWait实例,等待元素加载/可点击状态,提升稳定性 - CAPTCHA等待改为等待主页元素出现,无需固定60秒延迟
- 初始化
重置状态与容错处理:
- 每页爬取重新获取卡片列表,无需维护
selector_counter,避免状态混乱 - 增加异常捕获,处理单个卡片的爬取错误,不中断整体流程
- 每页爬取重新获取卡片列表,无需维护
文件写入优化:
- 指定编码为
utf-8,避免中文乱码 - 重命名输出文件为更清晰的
coursera_accomplishments.txt
- 指定编码为
内容的提问来源于stack exchange,提问作者wheeliefun
相关产品推荐
相关产品推荐

