使用Selenium爬取多个URL时出现IndexError,如何解决?
问题描述
使用Selenium WebDriver爬取多个博彩网站URL时,单独运行单个URL的爬取逻辑正常,但同时运行所有逻辑时抛出IndexError: list index out of range,报错发生在爬取getsbet网站的dictionar['1'] = odds_gb[x].text行。
代码片段
url_fortuna = "https://efortuna.ro/" url_betfair = 'https://www.betfair.ro/sport/' url_mozzart = "https://www.mozzartbet.ro/ro#/betting" driver = webdriver.Chrome() wait = WebDriverWait(driver, 20) driver.get(url_mozzart) matches_mz = driver.find_elements( By.CLASS_NAME, value='pairs' ) odds_mz = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR,"span[class='odd-font betting-regular-match font-18']"))) for event_mz in matches_mz: dictionar = {'casa': '', 'nume': '', '1': '', 'x': '', '2': ''} dictionar['casa'] = 'mozzart' match_mz = event_mz.text dictionar['nume'] = match_mz dictionar['1'] = odds_mz[x].text dictionar['x'] = odds_mz[x+1].text dictionar['2'] = odds_mz[x+2].text x= x+3 evenimente.append(dictionar) driver.get(url_unibet) matches_ub = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR,"div[data-test-name='scoreboard']"))) odds_ub = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR,"div[data-test-name='outcomeBet']"))) for event_ub in matches_ub: dictionar = {'casa': '', 'nume': '', '1': '', 'x': '', '2': ''} dictionar['casa'] = 'unibet' match_ub = event_ub.text dictionar['nume'] = match_ub dictionar['1'] = odds_ub[x].text dictionar['x'] = odds_ub[x+1].text dictionar['2'] = odds_ub[x+2].text x= x+3 evenimente.append(dictionar) driver.get(url_getsbet) WebDriverWait(driver, 20).until(EC.element_to_be_clickable((By.XPATH, "//button[@class='cookie_consent_btn']"))).click() WebDriverWait(driver, 20).until(EC.frame_to_be_available_and_switch_to_it((By.XPATH,"//iframe[@id='SportsBookIframe']"))) events_gb = WebDriverWait(driver, 20).until( EC.visibility_of_all_elements_located((By.XPATH, "//span[@class='Details__Participants']"))) odds_gb = WebDriverWait(driver, 20).until( EC.visibility_of_all_elements_located((By.XPATH, "//div[@class='OM-ValueChanger']"))) for event_gb in events_gb: dictionar = {'casa': '', 'nume': '', '1': '', 'x': '', '2': ''} dictionar['casa'] = 'getsbet' match_gb = event_gb.text dictionar['nume'] = match_gb dictionar['1'] = odds_gb[x].text dictionar['x'] = odds_gb[x+1].text dictionar['2'] = odds_gb[x+2].text x= x+3 evenimente.append(dictionar) driver.quit()
报错信息
Traceback (most recent call last): File "C:\Users\flr_1\PycharmProjects\scrapper\main.py", line 81, in <module> dictionar['1'] = odds_gb[x].text IndexError: list index out of range
解决方案
1. 重置索引变量x
问题核心是全局变量x在爬取第一个网站后持续累加,到第三个网站时,x的数值已经远超过odds_gb列表的长度,导致索引越界。每个网站的爬取逻辑需独立初始化索引:
# 爬取mozzart前初始化x x = 0 for event_mz in matches_mz: # 原有逻辑不变 x += 3 # 爬取unibet前重置x x = 0 for event_ub in matches_ub: # 原有逻辑不变 x += 3 # 爬取getsbet前重置x x = 0 for event_gb in events_gb: # 原有逻辑不变 x += 3
2. 修复iframe上下文问题
爬取getsbet时切换到了iframe,爬取完成后需切回主文档,避免后续操作(若有)出现元素定位失效:
# 爬取完getsbet后切回主文档 driver.switch_to.default_content()
3. 优化索引逻辑(可选)
用enumerate替代手动维护索引x,代码更简洁且不易出错:
# 以mozzart为例 for idx, event_mz in enumerate(matches_mz): dictionar = { 'casa': 'mozzart', 'nume': event_mz.text, '1': odds_mz[idx*3].text, 'x': odds_mz[idx*3+1].text, '2': odds_mz[idx*3+2].text } evenimente.append(dictionar)
4. 添加元素数量校验
确保每个网站的赛事数量乘以3等于赔率数量,避免页面结构变化导致的索引越界:
# 爬取mozzart时添加校验 if len(odds_mz) != len(matches_mz)*3: print(f"Mozzart: 赛事数量({len(matches_mz)})与赔率数量({len(odds_mz)})不匹配") # 可选择跳过当前网站或处理异常
内容的提问来源于stack exchange,提问作者Florentin-Alexandru Iosif
相关产品推荐
相关产品推荐

