如何用Python+Xpath提取网页中‘Nb B’行的所有数据?
Got it, let's tweak your script so it only pulls data from the 'Nb B' column instead of grabbing all table cells. Here's how to do it properly:
First, the issue with your original code is that your XPath is targeting every <td> with the odds class, which grabs all the odds values, return rates, and empty cells across the entire table. We need to be precise and target only the column corresponding to the 'Nb B' header.
Here's the updated script, with explanations of the key changes:
#!/usr/bin/python3 # -*- coding: utf-8 -*- from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By options = Options() options.headless = True driver = webdriver.Chrome(options=options) driver.get('https://www.coteur.com/cotes-foot.php') # Keep your existing code for game URLs if you need it url_links = [] for i in driver.find_elements(By.XPATH, '//a[contains(@href, "match/cotes-")]'): url_links.append(i.get_attribute('href')) print(len(url_links), '\n') # Extract only data from the 'Nb B' column nb_bookies = [] # Step 1: Locate the 'Nb B' header to find its column position nb_b_header = driver.find_element(By.XPATH, '//th[text()="Nb B"]') # Step 2: Calculate the 1-based column index (count previous headers +1) column_index = len(nb_b_header.find_elements(By.XPATH, './preceding-sibling::th')) + 1 # Step 3: Grab all cells in this column from the data table nb_b_cells = driver.find_elements(By.XPATH, f'//table[@id="table_cotes"]//tr/td[{column_index}]') # Iterate through cells, clean up text, and collect non-empty values for cell in nb_b_cells: cleaned_text = cell.text.strip() if cleaned_text: nb_bookies.append(cleaned_text) print(nb_bookies) # Don't forget to close the driver to free up resources driver.quit()
Key improvements:
- Precise targeting: Instead of broad
<td>selection, we first find the 'Nb B' header to get its column position. This makes the code resilient if the table's column order changes later. - Modern Selenium syntax: Switched to
By.XPATHinstead of the deprecatedfind_elements_by_xpathmethod, which is the recommended approach for Selenium 4 and above. - Clean data: Added
strip()to remove extra whitespace and a check to skip empty cells, so your final list only contains valid 'Nb B' values. - Resource management: Added
driver.quit()to properly close the browser session after scraping.
When you run this, you'll get a list of only the numbers from the 'Nb B' column, which represents the number of bookmakers for each soccer match.
内容的提问来源于stack exchange,提问作者ahmedaao
相关产品推荐
相关产品推荐

