如何通过Python匹配报告名,提取odd/even类元素首TD数值?
Got it, let's work through this problem together — I’ve dealt with similar web scraping tasks before, so here’s a straightforward, reliable solution using Selenium in Python that should get you what you need.
First, let’s clarify what we’re targeting: we need all table rows with class="odd" or class="even". For each row, we’ll check the 3rd <td> (your report name column) against your target string, and if there’s a match, pull the value from the 1st <td>.
Here’s a working implementation using modern Selenium syntax (since the old find_element_by_xpath method is deprecated now):
from selenium import webdriver from selenium.webdriver.common.by import By # Your target report name (the string you want to match) target_report = "Name of Report 2 (scrubbed)" # Initialize your browser driver (swap with Firefox/Edge if needed) driver = webdriver.Chrome() driver.get("your_target_page_url_here") # Grab all rows with either 'odd' or 'even' class matching_rows = driver.find_elements(By.XPATH, "//tr[@class='odd' or @class='even']") # Iterate through each row to find your match result_value = None for row in matching_rows: # Get the 3rd TD's text (XPath uses 1-based indexing, so [3] is the third column) report_name = row.find_element(By.XPATH, "./td[3]").text.strip() # Compare with your target string if report_name == target_report: # Extract the first TD's numeric value result_value = row.find_element(By.XPATH, "./td[1]").text.strip() break # Exit the loop once we find the match to save time # Use the result if result_value: print(f"Found matching value: {result_value}") else: print("No matching report name found on the page.") # Clean up the driver driver.quit()
Key Details to Note:
- Targeting Rows Correctly: The XPath
//tr[@class='odd' or @class='even']ensures we only pull rows with either of those two classes — no extra rows included. - Relative XPath for TDs: Using
./td[3]and./td[1]means we’re only looking within the current row, not the entire page. This prevents accidental matches from unrelated table cells. - Text Stripping: Adding
.strip()removes hidden whitespace (like line breaks or leading/trailing spaces) that might cause mismatches even if the text looks identical at first glance. - Waits (If Needed): If the page loads slowly, add a wait to make sure elements are present before scraping. For example:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # Wait up to 10 seconds for rows to load WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.XPATH, "//tr[@class='odd' or @class='even']")) )
Troubleshooting Tips:
- If matches aren’t showing up, print the
report_namevariable to check if there are hidden characters (like non-breaking spaces) in the page’s text that don’t match your target string. - If the rows are inside a specific table, narrow down the XPath to that table (e.g.,
//table[@id='report-table']//tr[@class='odd' or @class='even']) to avoid scraping rows from other tables on the page.
This approach should reliably find your matching row and extract the value you need. Let me know if you hit any edge cases that need tweaking!
内容的提问来源于stack exchange,提问作者pxl

