如何从页面中两个同class的表格提取仅值数据
Got it, let's work through this problem together. You're trying to extract just the text values (no HTML tags) from two tables that share the same prop-list class, and you've hit snags with both BeautifulSoup's find_all and Selenium XPath. Here are two straightforward solutions tailored to your use case:
If the page content loads fully without needing JavaScript interactions, this is the faster approach. We'll target all prop-list tables, then pull out only the text from their value cells:
import requests from bs4 import BeautifulSoup # Fetch the page content url = "https://www.redbook.com.au/cars/details/2019-honda-civic-50-years-edition-auto-my19/SPOT-ITM-524208/" response = requests.get(url) soup = BeautifulSoup(response.text, "html.parser") # Grab all tables with the target class tables = soup.find_all("table", class_="prop-list") # Extract values from each table separately for table_num, table in enumerate(tables, 1): print(f"--- Table {table_num} Values ---") # Find every 'value' cell in the current table value_cells = table.find_all("td", class_="value") for cell in value_cells: # Get clean text (strip removes extra whitespace) print(cell.get_text(strip=True))
Output you'll get:
--- Table 1 Values --- Alloy 215/55 R16 16x7.0 215/55 R16 16x7.0 --- Table 2 Values --- Rack and Pinion
Your tables are inside a hidden div (display: none), so you might need to first expand the collapsible section before extracting values. Here's how to fix your XPath approach to get pure text:
from selenium import webdriver from selenium.webdriver.common.by import By driver = webdriver.Chrome() driver.get(url) # Optional: Expand the collapsible panel (since tables are hidden by default) # Adjust the index if your tables are in a different collapsible section collapsible_headers = driver.find_elements(By.CLASS_NAME, "bh_collapsible-header") collapsible_headers[4].click() # Index starts at 0; matches your original XPath's section # Extract values from the first table print("--- Table 1 Values ---") table1_values = driver.find_elements(By.XPATH, "//table[@class='prop-list'][1]//td[@class='value']") for elem in table1_values: print(elem.text) # elem.text automatically gives you just the text, no tags # Extract values from the second table print("\n--- Table 2 Values ---") table2_values = driver.find_elements(By.XPATH, "//table[@class='prop-list'][2]//td[@class='value']") for elem in table2_values: print(elem.text) driver.quit()
Why your original XPath didn't work:
You were targeting a single td element instead of all value cells in the table. Using find_elements (plural) with the path //table[@class='prop-list'][1]//td[@class='value'] grabs every value cell in the first table, and elem.text strips out all HTML tags to give you just the content you need.
内容的提问来源于stack exchange,提问作者user11669928

