You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Selenium的URL标题脚本优化需求:TXT读URL、超时等待、自定义序号

Optimized URL Title Extractor with Custom Requirements

Below is the modified code that implements all three requested optimizations, along with detailed explanations for each change:

from pyvirtualdisplay import Display
from time import sleep
import sys
from selenium import webdriver
from selenium.common.exceptions import TimeoutException, WebDriverException
from selenium.webdriver.firefox.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# Fix encoding (kept for compatibility with older Python versions)
if hasattr(sys, 'setdefaultencoding'):
    reload(sys)
    sys.setdefaultencoding('utf-8')

def read_urls_from_file(file_path):
    """Read URLs from a text file, one per line, skipping empty lines"""
    try:
        with open(file_path, 'r', encoding='utf-8') as f:
            urls = [line.strip() for line in f if line.strip()]
        return urls
    except FileNotFoundError:
        print(f"Error: The file {file_path} was not found.")
        sys.exit(1)

def wait_for_title_change(driver, initial_title="Loading...", timeout=10):
    """Wait for the page title to change from the initial 'Loading...' state"""
    try:
        WebDriverWait(driver, timeout).until(
            EC.title_is_not(initial_title)
        )
        return True
    except TimeoutException:
        return False

if __name__ == "__main__":
    # Configuration - adjust these values as needed
    URL_FILE = "urls.txt"  # Path to your URL list file
    STARTING_INDEX = 50    # Custom starting number for output
    DISPLAY_SIZE = (800, 600)
    PAGE_LOAD_TIMEOUT = 60
    TITLE_WAIT_TIMEOUT = 10

    # Initialize virtual display
    display = Display(visible=0, size=DISPLAY_SIZE)
    display.start()

    # Configure Firefox driver
    options = Options()
    # Optional: Uncomment below to enable headless mode (no GUI)
    # options.add_argument("--headless")
    driver = webdriver.Firefox(executable_path='/usr/local/lib/geckodriver/geckodriver')
    driver.set_page_load_timeout(PAGE_LOAD_TIMEOUT)

    # Load URLs from text file
    urls = read_urls_from_file(URL_FILE)
    current_index = STARTING_INDEX

    for url in urls:
        print(f"Processing URL: {url}")
        try:
            driver.get(url)
            # Wait for title to exit "Loading..." state
            title_loaded = wait_for_title_change(driver, timeout=TITLE_WAIT_TIMEOUT)
            if title_loaded:
                print(f"{driver.title}, {current_index}")
            else:
                print(f"Title didn't load., {current_index}")
        except TimeoutException:
            print(f"Timeout loading URL., {current_index}")
        except WebDriverException as e:
            print(f"Error accessing {url}: {str(e)}, {current_index}")
        finally:
            current_index += 1
            # Small delay to avoid overwhelming target servers
            sleep(1)

    # Cleanup resources
    driver.quit()
    display.stop()

Key Optimizations Breakdown

1. Read URLs from Text File

  • Created a reusable read_urls_from_file function that loads URLs from a specified .txt file (one URL per line).
  • Automatically skips empty lines to avoid processing invalid entries.
  • Includes error handling for missing files to make the script more robust.

2. Wait for Title to Exit "Loading..." State

  • Used Selenium's WebDriverWait with EC.title_is_not to efficiently wait for the title to change (instead of fixed sleep times).
  • Implemented a 10-second timeout; if the title stays "Loading..." beyond that, it outputs the required message and moves to the next URL.

3. Custom Starting Index with Increment

  • Added a STARTING_INDEX variable to set your desired initial number (e.g., 50).
  • The current_index variable increments by 1 after processing each URL, ensuring consistent numbering across all outputs.

Extra Quality-of-Life Improvements

  • Added general WebDriverException handling to catch issues like invalid URLs or connection errors.
  • Centralized configuration variables at the top for easy adjustments.
  • Included optional headless mode for Firefox (useful if you don't need a virtual display in your environment).
  • Added a small delay between URL requests to be respectful of target website servers.

内容的提问来源于stack exchange,提问作者WeekSky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:50:16