如何在C#控制台程序中无第三方库修改HTML元素并填充XML数据?
Got it, since you can't use third-party libraries and need a native C# console solution, here's a straightforward approach using built-in XML parsing and regex for HTML manipulation. This works perfectly for your specific scenario where you map XML node keys to span IDs and replace their inner content.
Step 1: Parse XML into a Key-Value Dictionary
First, we'll load your XML file and convert its nodes into a dictionary for quick lookups. This assumes your XML has top-level elements where the element name is the "key" matching your span ID, and the element's value is what you want to insert into the span.
using System; using System.Collections.Generic; using System.Xml.Linq; // Load XML and build a key-value map var xmlDoc = XDocument.Load("your-data.xml"); var xmlKeyValuePairs = new Dictionary<string, string>(StringComparer.OrdinalIgnoreCase); foreach (var element in xmlDoc.Root.Elements()) { xmlKeyValuePairs[element.Name.LocalName] = element.Value.Trim(); }
Step 2: Modify HTML Using Regex
Next, we'll read your HTML file, use regex to find all <span> elements with an id attribute, and replace their content with the corresponding XML value (if a match exists).
Note: Regex isn't ideal for full HTML parsing, but since you're targeting specific, flat <span> elements (no nested spans), it's reliable for this controlled use case.
using System.IO; using System.Text.RegularExpressions; // Read original HTML content string originalHtml = File.ReadAllText("input.html"); // Regex pattern to match spans with ID attributes (captures ID and inner content) string spanRegexPattern = @"<span\s+id=""(?<spanId>[^""]+)""[^>]*>(?<innerContent>.*?)<\/span>"; // Replace matching spans with XML values string modifiedHtml = Regex.Replace( originalHtml, spanRegexPattern, match => { string spanId = match.Groups["spanId"].Value; if (xmlKeyValuePairs.TryGetValue(spanId, out string replacementText)) { // Return the span with updated content return $"<span id=\"{spanId}\">{replacementText}</span>"; } // If no XML match exists, leave the span unchanged return match.Value; }, RegexOptions.IgnoreCase | RegexOptions.Singleline ); // Save the modified HTML to a file File.WriteAllText("output.html", modifiedHtml);
Key Considerations
- Quotation Marks: This regex assumes your span
idattributes use double quotes. If your HTML uses single quotes, adjust the pattern toid='(?<spanId>[^']+)'. - Nested Spans: If your HTML has nested
<span>elements, this regex will break. But since you're working with a controlled project setup, this shouldn't be an issue. - XML Structure: Adjust the XML parsing code if your nodes aren't top-level (e.g., if they're under a specific parent element like
<data>).
内容的提问来源于stack exchange,提问作者Romans

