使用Selenium抓取IMDb watchlist仅能获取前100条内容如何解决
问题原因
- 执行逻辑顺序错误:代码将内容提取步骤放在了点击「加载更多」按钮之前,最后一次点击加载按钮后新增的内容,会因为下一轮循环找不到已消失的「加载更多」按钮抛出异常,直接终止程序,没有被提取到。
- 缺少新内容加载完成的判断:点击加载按钮后仅固定休眠5秒,没有验证新内容是否已经完全插入到
lister-list mode-detail节点下,可能出现休眠结束后新内容还未渲染完成的情况,导致获取的innerHTML不完整。 - 重复提取冗余内容:每次循环都解析整个列表的全部内容,重复输出已经打印过的条目,也会干扰对最终总数量的判断。
修正方案
调整代码执行顺序,先加载新内容再提取数据,添加加载完成判断逻辑,避免固定休眠的不确定性,示例代码如下:
from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from bs4 import BeautifulSoup import time URL = "https://www.imdb.com/user/ur130279232/watchlist" driver = webdriver.Chrome() wait = WebDriverWait(driver, 20) driver.get(URL) last_item_count = 0 all_items = [] while True: # 等待当前列表稳定 time.sleep(2) # 提取当前全量列表数据 watchlist = driver.find_element(By.XPATH, "//div[@class='lister-list mode-detail']") soup = BeautifulSoup(watchlist.get_attribute('innerHTML'), 'html.parser') current_items = soup.find_all('h3', class_ ='lister-item-header') current_count = len(current_items) # 条目数量无增长说明全部加载完成,终止循环 if current_count == last_item_count: break # 记录新增条目,避免重复打印 new_items = current_items[last_item_count:current_count] print(f'本次新增条数:{len(new_items)},当前总条数:{current_count}') for item in new_items: item_name = item.find('a').contents[0] all_items.append(item_name) print(item_name) last_item_count = current_count # 尝试点击加载更多按钮 try: load_more_btn = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@class='load-more']"))) load_more_btn.click() # 等待新条目加载完成 wait.until(lambda d: len(d.find_elements(By.XPATH, "//h3[@class='lister-item-header']")) > current_count) except: break # 输出最终统计 print(f'全部加载完成,总条数:{len(all_items)}') driver.quit()
内容的提问来源于stack exchange,提问作者Lus_Bus
相关产品推荐
相关产品推荐

