如何编写函数提取XML标签<>中的所有字符串并返回列表?
To solve this problem, we can use regular expressions to efficiently match and extract the tag names between < and > (ignoring closing tags like </TagName>). Here's a straightforward implementation tailored to your needs:
Step 1: Add Required Namespaces
First, include these namespaces at the top of your C# file to access the necessary classes:
using System.Collections.Generic; using System.Text.RegularExpressions;
Step 2: Implement the getTypes Function
This function uses a regex pattern to find all opening tag names, captures them, and returns them as a List<string>:
public List<string> getTypes(string input) { List<string> tagNames = new List<string>(); // Regex pattern to match opening tags (e.g., <FirstName>) // Captures content inside <> that doesn't start with a slash (avoids closing tags) string pattern = @"<([^/>]+)>"; MatchCollection matches = Regex.Matches(input, pattern); foreach (Match match in matches) { // Extract the captured tag name (group 1) and add to the list if (match.Groups.Count > 1) { string cleanTagName = match.Groups[1].Value.Trim(); // Skip empty strings in case tags have extra whitespace if (!string.IsNullOrEmpty(cleanTagName)) { tagNames.Add(cleanTagName); } } } return tagNames; }
Regex Pattern Breakdown
<: Matches the opening angle bracket.([^/>]+): Captures one or more characters that are not/or>(this skips closing tags like</FirstName>).>: Matches the closing angle bracket.
Step 3: Use the Function with Your Example
Here's how to integrate this function into your code as you described:
string s = "<FirstName> John </FirstName> <LastName> Snow </LastName>"; List<string> container = getTypes(s); foreach(string item in container) { Console.WriteLine(item); }
Expected Output
FirstName LastName
Optional: Get Unique Tag Names
If you want to avoid duplicate tag names (e.g., if the input has multiple instances of the same tag), use a HashSet to store unique values before converting back to a list:
public List<string> getUniqueTypes(string input) { HashSet<string> uniqueTags = new HashSet<string>(); string pattern = @"<([^/>]+)>"; MatchCollection matches = Regex.Matches(input, pattern); foreach (Match match in matches) { if (match.Groups.Count > 1) { string cleanTagName = match.Groups[1].Value.Trim(); if (!string.IsNullOrEmpty(cleanTagName)) { uniqueTags.Add(cleanTagName); } } } return uniqueTags.ToList(); }
Quick Note
This solution works great for simple tag structures. For complex HTML/XML (like nested tags or malformed markup), consider using a dedicated parser (such as XmlDocument or HtmlAgilityPack) instead of regex—regex can struggle with edge cases in full markup languages.
内容的提问来源于stack exchange,提问作者Gashio Lee

