Selenium爬取命运2战绩遇TimeoutException但仍打印日期的问题求助
解决Selenium爬取《命运2》战绩时的TimeoutException问题
我是Selenium新手,目标是爬取《命运2》中我的每场游戏胜利场次的战绩及时间。以下是我的代码:
driver = webdriver.Chrome() url = "https://destinytracker.com/destiny-2/profile/psn/4611686018440125811/matches?mode=crucible" driver.get(url) wait = WebDriverWait(driver, 30) ignored_exceptions=(NoSuchElementException,StaleElementReferenceException) crucible_content = wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, "div.trn-gamereport-list.trn-gamereport-list--compact"))) game_reports = crucible_content.find_elements(By.CLASS_NAME, "trn-gamereport-list__group") for game_report in game_reports: group_entry = wait.until(EC.presence_of_element_located((By.CLASS_NAME, "trn-gamereport-list__group-entries"))) win_match = group_entry.find_elements(By.CLASS_NAME,"trn-match-row--outcome-win") driver.execute_script("arguments[0].scrollIntoView();", win_match[0]) lose_match = group_entry.find_elements(By.CLASS_NAME, "trn-match-row--outcome-loss") for win_element in win_match: #try: #sleep(10) #driver.execute_script("arguments[0].scrollIntoView();", win_element) win_left = wait.until(EC.presence_of_element_located((By.CLASS_NAME, "trn-match-row__section--left"))) driver.execute_script("arguments[0].click();", win_left) print("reached here") date_time = wait.until(EC.presence_of_element_located((By.XPATH, "//div[@class='info']"))) date_time = date_time.text date_time = date_time.split(",") date_time = date_time[0] match_roster = wait.until(EC.presence_of_element_located((By.CLASS_NAME,"match-rosters"))) team_alpha = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "div.match-roster.alpha"))) team_bravo = wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "div.match-roster.bravo"))) bravo_match_roster_entries = team_bravo.find_element(By.CLASS_NAME, "roster-entries") alpha_match_roster_entries = team_alpha.find_element(By.CLASS_NAME,"roster-entries") name = wait.until(EC.presence_of_element_located((By.CLASS_NAME, "router-link-active"))) entry_bravo = bravo_match_roster_entries.find_elements(By.CLASS_NAME,"entry") entry_alpha = alpha_match_roster_entries.find_elements(By.CLASS_NAME,"entry") print(date_time)
运行后输出如下:
reached here 7/12/2023 reached here Traceback (most recent call last): File "/Users/victoruduma/Documents/python web scraping part 2 .py", line 121, in <module> crucible_dataset() File "/Users/victoruduma/Documents/python web scraping part 2 .py", line 55, in crucible_dataset date_time = wait.until(EC.presence_of_element_located((By.XPATH, "//div[@class='info']"))) File "/Library/Frameworks/Python.framework/Versions/3.11/lib/python3.11/site-packages/selenium/webdriver/support/wait.py", line 95, in until raise TimeoutException(message, screen, stacktrace) selenium.common.exceptions.TimeoutException: Message:
我已为此困扰一周,需要解决这个循环遍历胜利场次时的超时问题。
问题根源
- 全局元素定位冲突:
//div[@class='info']是全局查找,第二次点击新比赛后,页面可能存在多个同名元素,或旧元素未完全销毁,导致等待超时。 - 错误的元素绑定:
win_left每次都找全局第一个匹配元素,而非当前循环的win_element子元素,导致点击对象错误。 - 未处理弹窗状态:点击比赛详情后未关闭弹窗,残留的弹窗会干扰后续元素定位。
修正后的代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoSuchElementException, StaleElementReferenceException def crucible_dataset(): driver = webdriver.Chrome() url = "https://destinytracker.com/destiny-2/profile/psn/4611686018440125811/matches?mode=crucible" driver.get(url) # 初始化等待时直接忽略指定异常 wait = WebDriverWait(driver, 30, ignored_exceptions=(NoSuchElementException, StaleElementReferenceException)) # 等待比赛列表加载完成 crucible_content = wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, "div.trn-gamereport-list.trn-gamereport-list--compact"))) game_reports = crucible_content.find_elements(By.CLASS_NAME, "trn-gamereport-list__group") for game_report in game_reports: # 从当前分组内查找比赛条目,而非全局查找 group_entry = game_report.find_element(By.CLASS_NAME, "trn-gamereport-list__group-entries") win_match = group_entry.find_elements(By.CLASS_NAME,"trn-match-row--outcome-win") for win_element in win_match: # 滚动到当前胜利场次中心位置,确保可点击 driver.execute_script("arguments[0].scrollIntoView({block: 'center'});", win_element) # 从当前胜利场次元素内查找点击区域,避免全局找错 win_left = win_element.find_element(By.CLASS_NAME, "trn-match-row__section--left") win_left.click() print("reached here") # 等待详情弹窗加载完成,缩小定位范围到弹窗容器内 match_detail = wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, "div.trn-modal__content"))) # 用相对路径查找当前弹窗内的时间元素 date_time = match_detail.find_element(By.XPATH, ".//div[@class='info']") date_time_text = date_time.text.split(",")[0] # 从当前弹窗内查找队伍信息 match_roster = match_detail.find_element(By.CLASS_NAME,"match-rosters") team_alpha = match_roster.find_element(By.CSS_SELECTOR, "div.match-roster.alpha") team_bravo = match_roster.find_element(By.CSS_SELECTOR, "div.match-roster.bravo") bravo_entries = team_bravo.find_element(By.CLASS_NAME, "roster-entries").find_elements(By.CLASS_NAME,"entry") alpha_entries = team_alpha.find_element(By.CLASS_NAME,"roster-entries").find_elements(By.CLASS_NAME,"entry") print(date_time_text) # 关闭当前详情弹窗,避免影响下一次操作 close_btn = wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, "button.trn-modal__close"))) close_btn.click() # 等待弹窗完全消失后再进入下一次循环 wait.until(EC.invisibility_of_element_located((By.CSS_SELECTOR, "div.trn-modal__content"))) crucible_dataset()
关键修正点
- 缩小定位范围:所有详情页元素都从当前弹窗容器内查找,避免全局定位冲突。
- 绑定当前循环元素:点击的是当前
win_element的子元素,确保每次处理的是目标场次。 - 添加弹窗关闭逻辑:处理完单场详情后关闭弹窗,保证页面状态干净。
- 优化滚动逻辑:用
scrollIntoView({block: 'center'})提升元素点击成功率。
内容的提问来源于stack exchange,提问作者internshiphopeful
相关产品推荐
相关产品推荐

