Selenium页面滚动及目标URL检测停止滚动的技术问题咨询
Hey there! Let's work through your two Selenium questions step by step:
1. 如何按固定像素(1000px)滚动页面?
Instead of scrolling straight to the bottom with window.scrollTo(0, document.body.scrollHeight), you can use the window.scrollBy() method—it scrolls relative to the page's current position. To scroll down 1000 pixels each time, just replace that scroll line with:
driver.execute_script("window.scrollBy(0, 1000);")
This command tells the browser to move 0 pixels horizontally and 1000 pixels vertically from its current spot, giving you controlled, incremental scrolling.
2. 找到目标URL后停止滚动的问题
Your current code only breaks out of the inner for loop when it finds the target URL, but the outer while loop keeps running. Here's how to fix this:
- Add a flag variable (like
found_target) to track if we've located the URL - Once the target is found, set the flag to
Trueand break theforloop - Check the flag in the
whileloop to exit entirely
I also spotted a couple small syntax errors in your original code (missing quotes around target_url, a missing closing parenthesis in urls.append()). Here's the corrected full code:
from selenium.webdriver.common.by import By import time urls = [] target_url = "https://example.com" # Added missing quotes pause_time = 0.5 found_target = False # Flag to track if target URL is found while not found_target: links = driver.find_elements(By.XPATH, '//div[contains(@data-urn, "urn:li:activity:")]') for link in links: current_href = link.get_attribute('href') if current_href != target_url: urls.append(current_href) # Fixed missing closing parenthesis else: found_target = True break # Exit the for loop once target is found if found_target: break # Exit the while loop immediately # Scroll down 1000 pixels driver.execute_script("window.scrollBy(0, 1000);") time.sleep(pause_time) # Prevent infinite loop if we reach the bottom without finding the target new_height = driver.execute_script("return document.body.scrollHeight") current_scroll_pos = driver.execute_script("return window.pageYOffset + window.innerHeight") if current_scroll_pos >= new_height: break
A quick extra tip: Storing current_href in a variable avoids calling get_attribute('href') twice, which makes the code a bit more efficient.
备注:内容来源于stack exchange,提问作者LaurenLai

