You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python+Xpath提取网页中‘Nb B’行的所有数据?

Got it, let's tweak your script so it only pulls data from the 'Nb B' column instead of grabbing all table cells. Here's how to do it properly:

First, the issue with your original code is that your XPath is targeting every <td> with the odds class, which grabs all the odds values, return rates, and empty cells across the entire table. We need to be precise and target only the column corresponding to the 'Nb B' header.

Here's the updated script, with explanations of the key changes:

#!/usr/bin/python3
# -*- coding: utf-8 -*-
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By

options = Options()
options.headless = True
driver = webdriver.Chrome(options=options)
driver.get('https://www.coteur.com/cotes-foot.php')

# Keep your existing code for game URLs if you need it
url_links = []
for i in driver.find_elements(By.XPATH, '//a[contains(@href, "match/cotes-")]'):
    url_links.append(i.get_attribute('href'))
print(len(url_links), '\n')

# Extract only data from the 'Nb B' column
nb_bookies = []
# Step 1: Locate the 'Nb B' header to find its column position
nb_b_header = driver.find_element(By.XPATH, '//th[text()="Nb B"]')
# Step 2: Calculate the 1-based column index (count previous headers +1)
column_index = len(nb_b_header.find_elements(By.XPATH, './preceding-sibling::th')) + 1
# Step 3: Grab all cells in this column from the data table
nb_b_cells = driver.find_elements(By.XPATH, f'//table[@id="table_cotes"]//tr/td[{column_index}]')

# Iterate through cells, clean up text, and collect non-empty values
for cell in nb_b_cells:
    cleaned_text = cell.text.strip()
    if cleaned_text:
        nb_bookies.append(cleaned_text)

print(nb_bookies)
# Don't forget to close the driver to free up resources
driver.quit()

Key improvements:

  • Precise targeting: Instead of broad <td> selection, we first find the 'Nb B' header to get its column position. This makes the code resilient if the table's column order changes later.
  • Modern Selenium syntax: Switched to By.XPATH instead of the deprecated find_elements_by_xpath method, which is the recommended approach for Selenium 4 and above.
  • Clean data: Added strip() to remove extra whitespace and a check to skip empty cells, so your final list only contains valid 'Nb B' values.
  • Resource management: Added driver.quit() to properly close the browser session after scraping.

When you run this, you'll get a list of only the numbers from the 'Nb B' column, which represents the number of bookmakers for each soccer match.

内容的提问来源于stack exchange,提问作者ahmedaao

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 11:52:47