Python Selenium爬取Likibu.com时最后一页翻页停止逻辑问题求助
Hey there! The issue with your script is that the » button still exists on the last page—it just has an href="#" instead of a valid page URL. Your current check if not driver.find_element_by_link_text('»') will never trigger a break because the element is still present in the DOM.
Let's adjust your logic to check the button's href attribute instead. Here's how you can fix it:
Step-by-Step Solution
- Locate the "next page" button (
») safely (usingfind_elementsto avoidNoSuchElementExceptionif the button ever disappears unexpectedly). - Check if the button's
hrefattribute equals#—if yes, we're on the last page, so exit the loop. - If it's a valid URL, click the button to proceed to the next page.
Updated Code Snippet
# Locate the next page button (using find_elements to handle cases where it might not exist) next_buttons = driver.find_elements_by_link_text('»') if not next_buttons: # No next button found, exit loop break next_button = next_buttons[0] # Check if the href is "#" (last page) if next_button.get_attribute('href') == '#': break # Click to go to next page next_button.click()
Alternative Check Using Parent Element's Class
Looking at your provided HTML snippet, the last page's » button is wrapped in an <li> with the active class. You could also check for that class as an extra safeguard:
next_button = driver.find_element_by_link_text('»') parent_li = next_button.find_element_by_xpath('..') if parent_li.get_attribute('class') == 'active' or next_button.get_attribute('href') == '#': break next_button.click()
Pro Tip: Add Wait Conditions
To avoid race conditions (where the script tries to interact with the button before it's fully loaded), add an explicit wait:
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Wait up to 10 seconds for the next button to be clickable wait = WebDriverWait(driver, 10) next_button = wait.until(EC.element_to_be_clickable((By.LINK_TEXT, '»'))) if next_button.get_attribute('href') == '#': break next_button.click()
This way, your script will reliably stop as soon as it hits the last page, without trying to click the non-functional # link.
内容的提问来源于stack exchange,提问作者Mohammed Benaou

