基于Selenium的URL标题脚本优化需求:TXT读URL、超时等待、自定义序号
Optimized URL Title Extractor with Custom Requirements
Below is the modified code that implements all three requested optimizations, along with detailed explanations for each change:
from pyvirtualdisplay import Display from time import sleep import sys from selenium import webdriver from selenium.common.exceptions import TimeoutException, WebDriverException from selenium.webdriver.firefox.options import Options from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Fix encoding (kept for compatibility with older Python versions) if hasattr(sys, 'setdefaultencoding'): reload(sys) sys.setdefaultencoding('utf-8') def read_urls_from_file(file_path): """Read URLs from a text file, one per line, skipping empty lines""" try: with open(file_path, 'r', encoding='utf-8') as f: urls = [line.strip() for line in f if line.strip()] return urls except FileNotFoundError: print(f"Error: The file {file_path} was not found.") sys.exit(1) def wait_for_title_change(driver, initial_title="Loading...", timeout=10): """Wait for the page title to change from the initial 'Loading...' state""" try: WebDriverWait(driver, timeout).until( EC.title_is_not(initial_title) ) return True except TimeoutException: return False if __name__ == "__main__": # Configuration - adjust these values as needed URL_FILE = "urls.txt" # Path to your URL list file STARTING_INDEX = 50 # Custom starting number for output DISPLAY_SIZE = (800, 600) PAGE_LOAD_TIMEOUT = 60 TITLE_WAIT_TIMEOUT = 10 # Initialize virtual display display = Display(visible=0, size=DISPLAY_SIZE) display.start() # Configure Firefox driver options = Options() # Optional: Uncomment below to enable headless mode (no GUI) # options.add_argument("--headless") driver = webdriver.Firefox(executable_path='/usr/local/lib/geckodriver/geckodriver') driver.set_page_load_timeout(PAGE_LOAD_TIMEOUT) # Load URLs from text file urls = read_urls_from_file(URL_FILE) current_index = STARTING_INDEX for url in urls: print(f"Processing URL: {url}") try: driver.get(url) # Wait for title to exit "Loading..." state title_loaded = wait_for_title_change(driver, timeout=TITLE_WAIT_TIMEOUT) if title_loaded: print(f"{driver.title}, {current_index}") else: print(f"Title didn't load., {current_index}") except TimeoutException: print(f"Timeout loading URL., {current_index}") except WebDriverException as e: print(f"Error accessing {url}: {str(e)}, {current_index}") finally: current_index += 1 # Small delay to avoid overwhelming target servers sleep(1) # Cleanup resources driver.quit() display.stop()
Key Optimizations Breakdown
1. Read URLs from Text File
- Created a reusable
read_urls_from_filefunction that loads URLs from a specified.txtfile (one URL per line). - Automatically skips empty lines to avoid processing invalid entries.
- Includes error handling for missing files to make the script more robust.
2. Wait for Title to Exit "Loading..." State
- Used Selenium's
WebDriverWaitwithEC.title_is_notto efficiently wait for the title to change (instead of fixed sleep times). - Implemented a 10-second timeout; if the title stays "Loading..." beyond that, it outputs the required message and moves to the next URL.
3. Custom Starting Index with Increment
- Added a
STARTING_INDEXvariable to set your desired initial number (e.g., 50). - The
current_indexvariable increments by 1 after processing each URL, ensuring consistent numbering across all outputs.
Extra Quality-of-Life Improvements
- Added general
WebDriverExceptionhandling to catch issues like invalid URLs or connection errors. - Centralized configuration variables at the top for easy adjustments.
- Included optional headless mode for Firefox (useful if you don't need a virtual display in your environment).
- Added a small delay between URL requests to be respectful of target website servers.
内容的提问来源于stack exchange,提问作者WeekSky
相关产品推荐
相关产品推荐

