如何在Python爬虫中到达目标页面(第530页)时终止循环?
问题分析与修复方案
你的代码陷入无限循环的核心原因是依赖URL结尾判断是否到达目标页的逻辑不可靠——点击“加载更多”后,页面URL可能不会按预期更新,或者跳转逻辑和你设想的不一致。
更稳妥的方式是添加页面计数变量,每次点击“加载更多”后计数递增,当计数达到目标页时直接终止循环。以下是具体修改方案:
修改后的完整代码(关键改动已标注)
import time from selenium.webdriver.common.by import By from selenium.webdriver.support.wait import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium import webdriver from webdriver_manager.chrome import ChromeDriverManager from tqdm import tqdm import pandas as pd driver = webdriver.Chrome(ChromeDriverManager().install()) def get_articles(last_page): url = "https://www.jimin.jp/news/?more=25" driver.get(url) aricle_xpath = "/html/body/div[2]/div/div[3]/div/div[1]/div[4]/div/div" wait = WebDriverWait(driver, 30) element = wait.until(EC.element_to_be_clickable((By.XPATH, aricle_xpath))) pre = -1 scrollHeight = 500 interval = 500 retry = 0 total_retry = 5 articles = [] current_page = 25 # 初始页面对应more=25 # 主循环改为用页面计数判断 while current_page < last_page: # 子循环:等待当前页内容加载完成 while pre != len(articles) and retry != total_retry: pre = len(articles) driver.execute_script(f"window.scrollTo(0, {scrollHeight});") scrollHeight += interval articles = driver.find_elements("xpath", aricle_xpath) time.sleep(1.5) retry += 1 pre = -1 retry = 0 # 提前检查是否到达目标页,避免多余点击 if current_page >= last_page: break more_button_xpath = "/html/body/div[2]/div/div[3]/div/div[1]/div[5]/div/a" # 等待按钮可点击,防止页面未加载完成报错 wait.until(EC.element_to_be_clickable((By.XPATH, more_button_xpath))).click() time.sleep(1.5) current_page += 25 # 每次加载更多,页面参数增加25 print(f"Collected articles: {len(articles)}, Current page: {current_page}", end="\r") # 提取所有文章链接 articles = driver.find_elements("xpath", aricle_xpath) link_xpath = "./div/a" urls = [] for article in articles: urls.append(article.find_element("xpath", link_xpath).get_attribute("href")) return urls def get_article_data(url): driver.get(url) time.sleep(1) date_xpath = "/html/body/div[2]/div/div[2]/div[3]/div[1]/div/div[1]/div" body_xpath = "/html/body/div[2]/div/div[2]/div[3]/div[2]" date = driver.find_element("xpath", date_xpath).text body = driver.find_element("xpath", body_xpath).text return date, body last_page = 530 urls = get_articles(last_page) data = [] for url in tqdm(urls): if url.startswith("https://www.jimin.jp/news/information/"): # 避免重复调用接口,优化性能 date, body = get_article_data(url) data.append((date, body)) time.sleep(1)
关键改动说明
- 替换循环判断逻辑:将依赖URL的判断改为
current_page计数,初始值匹配初始URL的more=25,每次点击加载更多后递增25,当计数达到last_page时直接跳出循环,彻底解决无限循环问题。 - 添加提前检查:在点击“加载更多”前判断是否已到达目标页,避免多余的无效点击。
- 优化元素等待:用
WebDriverWait确保“加载更多”按钮可点击后再操作,减少页面加载延迟导致的报错。 - 修复重复调用问题:在提取文章数据时,避免重复调用
get_article_data,减少不必要的页面请求。
内容的提问来源于stack exchange,提问作者Julia
相关产品推荐
相关产品推荐

