基于Selenium与XPath的C#节点计数及赛事数据爬取问题
Absolutely, your approach is totally feasible—this is exactly how you'd scale from scraping a single event to batch-processing all live matches on the page. The key is to target the parent container of all match nodes, collect them into a list, then iterate through each one to extract data using relative XPath queries (far more reliable than absolute paths that break with minor page structure changes).
Step-by-Step Implementation
1. Target the Parent Container of All Matches
Looking at the HTML you provided, all live matches are wrapped in <li> elements inside a <ul class="events--list">. Instead of relying on the brittle absolute XPath //*[@id='livediv']/div/div[2]/ul/li[1]/ul/li[1], use a more stable selector to grab all match nodes:
// Get all live match nodes var matchNodes = driver.FindElements(By.XPath("//ul[@class='events--list']/li"));
2. Iterate Through Each Match Node
For each match node, use relative XPath (starting with .// to search within the current node) to extract the required data. This avoids hardcoding positions like [1] and works for every match in the list.
3. Add Robustness with Error Handling
Some matches might have missing elements (e.g., no live timer yet), so wrap element lookups in try/catch blocks or check for nulls to prevent crashes.
4. Clean Up Resources
Don't forget to quit the driver when done to free up system resources.
Full Modified Code
using OpenQA.Selenium; using OpenQA.Selenium.Chrome; using System; using System.Collections.Generic; using System.Text; class Program { static void Main(string[] args) { // Initialize ChromeDriver with automatic resource cleanup using (var driver = new ChromeDriver()) { driver.Navigate().GoToUrl("https://www.favorit.com.ua/uk/live/"); Console.OutputEncoding = Encoding.UTF8; // Ensure Ukrainian characters display correctly // Brief wait for page to load (replace with explicit waits for production code) System.Threading.Thread.Sleep(3000); // Get all live match nodes from the events list IList<IWebElement> matchNodes = driver.FindElements(By.XPath("//ul[@class='events--list']/li")); Console.WriteLine($"Found {matchNodes.Count} live matches:\n"); // Loop through each match to extract data foreach (var matchNode in matchNodes) { try { // Extract data using relative XPath (scoped to the current match node) string matchId = matchNode.FindElement(By.XPath(".//div[@class='event--line--position']/span")).Text; string team1 = matchNode.FindElement(By.XPath(".//div[@class='event--name two--name']/span[1]")).GetAttribute("innerHTML"); string team2 = matchNode.FindElement(By.XPath(".//div[@class='event--name two--name']/span[2]")).GetAttribute("innerHTML"); string team1Goals = matchNode.FindElement(By.XPath(".//div[@class='result--item result--item--cur']/div[1]")).Text; string team2Goals = matchNode.FindElement(By.XPath(".//div[@class='result--item result--item--cur']/div[2]")).Text; string matchTime = matchNode.FindElement(By.XPath(".//div[@class='event--timer']")).Text; // Print formatted match details Console.WriteLine($"Game ID: {matchId}"); Console.WriteLine($"{team1} {team1Goals} vs {team2} {team2Goals}"); Console.WriteLine($"Current Match Time: {matchTime}\n"); } catch (NoSuchElementException ex) { // Skip matches with incomplete data instead of crashing Console.WriteLine($"Skipping match: Missing element - {ex.Message}\n"); continue; } } Console.ReadLine(); } // Driver automatically quits when exiting the 'using' block } }
Key Improvements Over Your Original Code
- Relative XPath Queries:
.//ensures we only search within the current match node, making the code resilient to page structure changes. - Batch Processing: Collects all match nodes first, then iterates through them to extract data for every live event.
- Resource Management: Uses
usingblock to automatically dispose of the ChromeDriver when done. - Error Handling: Catches
NoSuchElementExceptionto skip matches with missing data instead of crashing the program. - Readable Output: Formats the output for better readability and debugging.
Optional Enhancements
- Explicit Waits: Replace
Thread.SleepwithWebDriverWaitto wait for specific elements to load, which is more efficient and reliable:var wait = new WebDriverWait(driver, TimeSpan.FromSeconds(10)); wait.Until(d => d.FindElement(By.XPath("//ul[@class='events--list']/li"))); - Data Storage: Instead of printing to console, save the extracted data to a CSV file, database, or directly to your online stats system.
内容的提问来源于stack exchange,提问作者The Vee

