You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Selenium与XPath的C#节点计数及赛事数据爬取问题

Batch Scraping Live Match Data from Favorit.com.ua

Absolutely, your approach is totally feasible—this is exactly how you'd scale from scraping a single event to batch-processing all live matches on the page. The key is to target the parent container of all match nodes, collect them into a list, then iterate through each one to extract data using relative XPath queries (far more reliable than absolute paths that break with minor page structure changes).

Step-by-Step Implementation

1. Target the Parent Container of All Matches

Looking at the HTML you provided, all live matches are wrapped in <li> elements inside a <ul class="events--list">. Instead of relying on the brittle absolute XPath //*[@id='livediv']/div/div[2]/ul/li[1]/ul/li[1], use a more stable selector to grab all match nodes:

// Get all live match nodes
var matchNodes = driver.FindElements(By.XPath("//ul[@class='events--list']/li"));

2. Iterate Through Each Match Node

For each match node, use relative XPath (starting with .// to search within the current node) to extract the required data. This avoids hardcoding positions like [1] and works for every match in the list.

3. Add Robustness with Error Handling

Some matches might have missing elements (e.g., no live timer yet), so wrap element lookups in try/catch blocks or check for nulls to prevent crashes.

4. Clean Up Resources

Don't forget to quit the driver when done to free up system resources.

Full Modified Code

using OpenQA.Selenium;
using OpenQA.Selenium.Chrome;
using System;
using System.Collections.Generic;
using System.Text;

class Program
{
    static void Main(string[] args)
    {
        // Initialize ChromeDriver with automatic resource cleanup
        using (var driver = new ChromeDriver())
        {
            driver.Navigate().GoToUrl("https://www.favorit.com.ua/uk/live/");
            Console.OutputEncoding = Encoding.UTF8; // Ensure Ukrainian characters display correctly

            // Brief wait for page to load (replace with explicit waits for production code)
            System.Threading.Thread.Sleep(3000);

            // Get all live match nodes from the events list
            IList<IWebElement> matchNodes = driver.FindElements(By.XPath("//ul[@class='events--list']/li"));

            Console.WriteLine($"Found {matchNodes.Count} live matches:\n");

            // Loop through each match to extract data
            foreach (var matchNode in matchNodes)
            {
                try
                {
                    // Extract data using relative XPath (scoped to the current match node)
                    string matchId = matchNode.FindElement(By.XPath(".//div[@class='event--line--position']/span")).Text;
                    string team1 = matchNode.FindElement(By.XPath(".//div[@class='event--name two--name']/span[1]")).GetAttribute("innerHTML");
                    string team2 = matchNode.FindElement(By.XPath(".//div[@class='event--name two--name']/span[2]")).GetAttribute("innerHTML");
                    string team1Goals = matchNode.FindElement(By.XPath(".//div[@class='result--item result--item--cur']/div[1]")).Text;
                    string team2Goals = matchNode.FindElement(By.XPath(".//div[@class='result--item result--item--cur']/div[2]")).Text;
                    string matchTime = matchNode.FindElement(By.XPath(".//div[@class='event--timer']")).Text;

                    // Print formatted match details
                    Console.WriteLine($"Game ID: {matchId}");
                    Console.WriteLine($"{team1} {team1Goals} vs {team2} {team2Goals}");
                    Console.WriteLine($"Current Match Time: {matchTime}\n");
                }
                catch (NoSuchElementException ex)
                {
                    // Skip matches with incomplete data instead of crashing
                    Console.WriteLine($"Skipping match: Missing element - {ex.Message}\n");
                    continue;
                }
            }

            Console.ReadLine();
        } // Driver automatically quits when exiting the 'using' block
    }
}

Key Improvements Over Your Original Code

  • Relative XPath Queries: .// ensures we only search within the current match node, making the code resilient to page structure changes.
  • Batch Processing: Collects all match nodes first, then iterates through them to extract data for every live event.
  • Resource Management: Uses using block to automatically dispose of the ChromeDriver when done.
  • Error Handling: Catches NoSuchElementException to skip matches with missing data instead of crashing the program.
  • Readable Output: Formats the output for better readability and debugging.

Optional Enhancements

  • Explicit Waits: Replace Thread.Sleep with WebDriverWait to wait for specific elements to load, which is more efficient and reliable:
    var wait = new WebDriverWait(driver, TimeSpan.FromSeconds(10));
    wait.Until(d => d.FindElement(By.XPath("//ul[@class='events--list']/li")));
    
  • Data Storage: Instead of printing to console, save the extracted data to a CSV file, database, or directly to your online stats system.

内容的提问来源于stack exchange,提问作者The Vee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:47:31