如何点击View Schedule按钮抓取.aspx站点的中队排班数据并指定日期?
Solution to Scrape VT-9 Schedule with Selenium (Supports Custom Dates)
Got it, let's work through this problem step by step. The core issue here is that clicking the "View Schedule" button loads the table via AJAX (no URL change), so you need to wait for the table to fully load after clicking instead of trying to scrape immediately. Plus, we can add code to set a custom date before triggering the button click.
Here's a revised, working version of your code with clear explanations:
Step-by-Step Code Implementation
First, make sure you're using the latest Selenium version (old methods like find_element_by_id are deprecated). Update it if needed with:
pip install --upgrade selenium
Now the full functional code:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import bs4 as bs # Configure your ChromeDriver path (match your OS and Chrome version) driver_path = "/users/base/Downloads/chromedriver" driver = webdriver.Chrome(executable_path=driver_path) try: # Open the target squadron schedule page url = "https://www.cnatra.navy.mil/scheds/schedule_data.aspx?sq=vt-9" driver.get(url) # -------------------------- # Optional: Set a custom date (use MM/DD/YYYY format to match the page's requirements) target_date = "10/15/2024" # Replace with your desired date date_input = WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.ID, "txtDate")) # Verify this ID via page inspection if needed ) date_input.clear() # Clear the default current date date_input.send_keys(target_date) # -------------------------- # Wait for the "View Schedule" button to be clickable, then click it view_button = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.ID, "btnViewSched")) ) view_button.click() # Critical: Wait for the schedule table to load (AJAX takes time to fetch data) schedule_table = WebDriverWait(driver, 15).until( EC.presence_of_element_located((By.TAG_NAME, "table")) # Use the table's specific ID if available for better accuracy ) # Option 1: Extract table text directly with Selenium print("Raw Schedule Data:\n", schedule_table.text) # Option 2: Parse with BeautifulSoup for structured data extraction page_source = driver.page_source soup = bs.BeautifulSoup(page_source, "lxml") table = soup.find("table") # You can now extract rows, columns, or specific cells from the parsed table print("\nStructured Table (BeautifulSoup):\n", table.prettify()) finally: # Ensure the browser closes even if an error occurs driver.quit()
Key Improvements Explained
- Explicit Waits: Instead of assuming elements load instantly, we use
WebDriverWaitto wait for elements to be present/clickable. This fixes race conditions where the script tries to interact with elements before they're ready. - Custom Date Support: We locate the date input field, clear the default value, and send your target date. Double-check the input field's ID via page inspection if
txtDatedoesn't work. - AJAX Handling: After clicking the button, we wait for the table to load before scraping. This ensures we capture the updated schedule data, not the empty initial page.
- Cleanup: The
finallyblock guarantees the browser closes properly, even if an error interrupts the script.
Troubleshooting Tips
- If the date input ID isn't
txtDate, right-click the date field on the page → Inspect to find the correct ID. - If the schedule table has a unique ID (instead of just
<table>), use that inEC.presence_of_element_locatedfor more precise targeting. - If you get a ChromeDriver version mismatch error, download the driver version that matches your installed Chrome browser.
内容的提问来源于stack exchange,提问作者exos
相关产品推荐
相关产品推荐

