基于正则表达式提取非规范HTML中特定MLA格式ID的C#方案
Hey there! Let's solve this problem of extracting those MLA-prefixed IDs from messy, non-standard HTML (since we can't use proper parsers like HtmlAgilityPack). I'll walk you through three different regex patterns in C#, each with its own use case, plus sample code and test HTML snippets.
First, let's look at some sample non-standard HTML we might be dealing with:
<!-- Sample non-well-formed HTML snippets --> <div id='MLA1234567' class="container">Some content</div> <p id="MLA98765432" data-value="123">Another paragraph</p> <span id=MLA112233445 >Unquoted attribute value</span> <div id="invalidMLA12345">Too short digits (only 5)</div> <div class="item" id=" MLA6543210987 ">Valid ID with extra spaces</div> <div id="MLA123">Way too short (3 digits)</div>
Our goal is to extract IDs that:
- Are inside an
idattribute - Start with
MLA - Followed by 7 or more digits (since "length超过6位" means 7+ digits)
Three Regex Patterns & C# Implementations
Each pattern targets different levels of HTML messiness. All implementations will collect matches into a List<string>.
1. Strict Quoted Attribute Pattern
Best for most cases where id values are wrapped in single or double quotes. This is the most precise and least likely to false-positive.
Regex Pattern:
id=["']MLA(\d{7,})["']
C# Code:
using System.Collections.Generic; using System.Text.RegularExpressions; public static List<string> ExtractStrictMlaIds(string htmlText) { var mlaIds = new List<string>(); const string pattern = @"id=["']MLA(\d{7,})["']"; foreach (Match match in Regex.Matches(htmlText, pattern, RegexOptions.IgnoreCase)) { // Reconstruct the full ID by combining MLA + captured digits mlaIds.Add($"MLA{match.Groups[1].Value}"); } return mlaIds; }
2. Unquoted Attribute Support Pattern
Handles non-standard HTML where id values aren't wrapped in quotes (e.g., id=MLA12345678). Uses a backreference to ensure consistent quoting (or no quotes at all).
Regex Pattern:
id\s*=\s*(["']?)MLA(\d{7,})\1
C# Code:
public static List<string> ExtractMlaIdsWithUnquotedSupport(string htmlText) { var mlaIds = new List<string>(); const string pattern = @"id\s*=\s*(["']?)MLA(\d{7,})\1"; foreach (Match match in Regex.Matches(htmlText, pattern, RegexOptions.IgnoreCase)) { mlaIds.Add($"MLA{match.Groups[2].Value}"); } return mlaIds; }
3. Whitespace-Tolerant Pattern
For extreme cases where id values have extra spaces (e.g., id=" MLA12345678 "). This pattern ignores whitespace inside the attribute value wrapper.
Regex Pattern:
id\s*=\s*(["']?)\s*MLA(\d{7,})\s*\1
C# Code:
public static List<string> ExtractMlaIdsWithWhitespaceTolerance(string htmlText) { var mlaIds = new List<string>(); const string pattern = @"id\s*=\s*(["']?)\s*MLA(\d{7,})\s*\1"; foreach (Match match in Regex.Matches(htmlText, pattern, RegexOptions.IgnoreCase)) { mlaIds.Add($"MLA{match.Groups[2].Value}"); } return mlaIds; }
Test Usage Example
Here's how you can test these methods with the sample HTML:
public static void Main() { string sampleHtml = @" <!-- Sample non-well-formed HTML snippets --> <div id='MLA1234567' class=""container"">Some content</div> <p id=""MLA98765432"" data-value=""123"">Another paragraph</p> <span id=MLA112233445 >Unquoted attribute value</span> <div id=""invalidMLA12345"">Too short digits (only 5)</div> <div class=""item"" id="" MLA6543210987 "" >Valid ID with extra spaces</div> <div id=""MLA123"">Way too short (3 digits)</div> "; // Test each method var strictIds = ExtractStrictMlaIds(sampleHtml); var unquotedIds = ExtractMlaIdsWithUnquotedSupport(sampleHtml); var whitespaceIds = ExtractMlaIdsWithWhitespaceTolerance(sampleHtml); System.Console.WriteLine("Strict Pattern Results:"); foreach (var id in strictIds) System.Console.WriteLine($"- {id}"); System.Console.WriteLine("\nUnquoted Support Results:"); foreach (var id in unquotedIds) System.Console.WriteLine($"- {id}"); System.Console.WriteLine("\nWhitespace-Tolerant Results:"); foreach (var id in whitespaceIds) System.Console.WriteLine($"- {id}"); }
Key Notes
- All patterns use
RegexOptions.IgnoreCaseto handle cases likeID="MLA1234567"(uppercase ID). - If you need to exclude IDs inside HTML comments (
<!-- ... -->), you can extend the regex to skip comment blocks, but that adds complexity. Let me know if you need that! - Regex isn't perfect for all HTML edge cases, but for non-standard text input, these patterns cover most common scenarios.
内容的提问来源于stack exchange,提问作者Mr John

