使用Python Selenium抓取网站中所有Tooltip(提示框)内容
Hey Jay, let's tackle this tooltip scraping problem for the Townsville Port schedule site. I've got two solid approaches for you—one using browser automation to mimic human interaction, and another that's more efficient by targeting the backend API directly.
方法一:用Selenium模拟悬停触发Tooltip
This is the most straightforward approach since it replicates exactly what a human would do: hover over each event element to make the tooltip appear, then extract its text.
步骤&代码示例
First, install the required dependencies:
pip install selenium webdriver-manager
Then use this script to scrape the tooltips:
from selenium import webdriver from selenium.webdriver.common.action_chains import ActionChains from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from webdriver_manager.chrome import ChromeDriverManager # Initialize Chrome browser (Firefox works too with minor tweaks) driver = webdriver.Chrome(ChromeDriverManager().install()) driver.get("https://schedule.townsville-port.com.au/") # Wait for the event elements to load on the page wait = WebDriverWait(driver, 10) events = wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "dhx_cal_event_line"))) collected_tooltips = [] for event in events: try: # Simulate hovering over the event to trigger the tooltip ActionChains(driver).move_to_element(event).perform() # Wait for the tooltip to become visible, then grab its text tooltip = wait.until(EC.visibility_of_element_located((By.CLASS_NAME, "dhtmlXTooltip"))) collected_tooltips.append(tooltip.text) # Optional: Print each tooltip as we collect it for debugging print(f"Tooltip text: {tooltip.text}") except Exception as e: print(f"Failed to process event: {str(e)}") continue # Clean up and close the browser driver.quit() # Output all collected tooltip texts print("\nAll collected tooltip texts:") for idx, text in enumerate(collected_tooltips, 1): print(f"{idx}. {text}")
方法二:直接抓取后端API数据(更高效)
Most calendar-style sites load event data via AJAX requests, which means the tooltip content is probably already available in a JSON response from the backend—no need to simulate a browser.
步骤
- Open your browser's DevTools (F12), switch to the Network tab, and filter by XHR/Fetch.
- Refresh the page or hover over an event—look for requests that return JSON data containing the
event_idvalues (like 55591 from your example). - Once you find the API endpoint, you can fetch the data directly with a simple HTTP request.
代码示例
import requests # Mimic a browser's request headers to avoid being blocked headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # Replace this with the actual API endpoint you find in DevTools api_url = "https://schedule.townsville-port.com.au/api/events" response = requests.get(api_url, headers=headers) events_data = response.json() collected_tooltips = [] for event in events_data: # You'll need to inspect the JSON structure to find the correct key for tooltip content # Common keys might be 'tooltip', 'description', or 'details' if "tooltip" in event: collected_tooltips.append(event["tooltip"]) print("All tooltip texts from API:") for text in collected_tooltips: print(text)
实用小贴士
- For Selenium: Use
WebDriverWaitinstead oftime.sleep()to avoid unnecessary delays and handle dynamic loading better. - For API scraping: Always check the request headers in DevTools to match what the browser sends (like
RefererorUser-Agent) to prevent being blocked. - If the site loads events by date ranges, adjust the API request parameters (like
start_dateorend_date) to fetch all relevant events.
内容的提问来源于stack exchange,提问作者Jay Haran

