如何使用Python Selenium处理意大利官方公报页面异常:部分页面无Visualizza按钮直接显示法律文本
Handling Dynamic Page States in Selenium for Gazzetta Ufficiale
Great question! This is a super common scenario when scraping dynamic sites—some pages need an extra click to reveal content, while others drop you straight into what you need. Here's a clean, reliable way to handle both cases:
The Core Idea
Instead of assuming the "Visualizza" button will always exist, we'll:
- Attempt to wait for and click the button (for pages that require it)
- Catch the
TimeoutExceptionif the button never appears (meaning we're already on the final content page) - Verify we're on the right page even when the button is missing, to avoid false positives
Modified Code Example
Here's how to adjust your script to handle both page types smoothly:
import time from selenium import webdriver from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException driver = webdriver.Chrome("/Users/bob/Documents/work/scraper/scrape_gu/chromedriver") def process_legal_page(driver, url): driver.get(url) try: # Wait up to 10 seconds for the "Visualizza" button to appear visualizza_button = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.XPATH, '//*[@id="corpo_export"]/div/input[1]')) ) visualizza_button.click() print("Clicked Visualizza button to access full text") # Wait for the final content page to load (update the selector to match your target element) WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "testo_atto")) # Replace with actual final page element ID ) except TimeoutException: # Button didn't appear—confirm we're already on the full text page try: WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "testo_atto")) # Use the same final page element ) print("Directly loaded full text page, no button needed") except TimeoutException: # Neither button nor content found—handle this error gracefully print(f"Error: Failed to load content for URL {url}") raise # Re-raise if you want to stop execution, or handle as needed # Test with both page types process_legal_page(driver, "https://www.gazzettaufficiale.it/atto/vediMenuHTML?atto.dataPubblicazioneGazzetta=2021-01-02&atto.codiceRedazionale=20A07300&tipoSerie=serie_generale&tipoVigenza=originario") time.sleep(5) process_legal_page(driver, "https://www.gazzettaufficiale.it/atto/vediMenuHTML?atto.dataPubblicazioneGazzetta=2021-01-02&atto.codiceRedazionale=20A07249&tipoSerie=serie_generale&tipoVigenza=originario") time.sleep(5) driver.quit()
Key Improvements
- Try-Except Block: Prevents your script from crashing when the button is missing by catching the
TimeoutException - Content Verification: Ensures you're actually on the page with the legal text even when the button isn't present (use your browser's dev tools to find a unique ID/class for the text container)
- Modular Function: Wraps the page logic into a reusable function, making it easy to drop into your existing loop
- Explicit Waits: Replaces arbitrary
time.sleep()calls with waits tied to actual page elements, making your script faster and more reliable
Quick Tips
- Inspect the Final Page: Use your browser's developer tools to find a unique identifier (like
testo_attoin the example) for the legal text area—this will make your content check more robust - Adjust Timeouts: Tweak the 10-second wait times based on how fast the site loads for you
- Handle Edge Cases: If you encounter other rare page states, add additional checks or exception handlers to cover them
Happy scraping, and belated happy 2022! 😊
内容的提问来源于stack exchange,提问作者Robert Alexander
相关产品推荐
相关产品推荐

