You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于正则表达式提取非规范HTML中特定MLA格式ID的C#方案

Hey there! Let's solve this problem of extracting those MLA-prefixed IDs from messy, non-standard HTML (since we can't use proper parsers like HtmlAgilityPack). I'll walk you through three different regex patterns in C#, each with its own use case, plus sample code and test HTML snippets.

Extracting MLA-prefixed IDs from Non-Standard HTML with C# Regex

First, let's look at some sample non-standard HTML we might be dealing with:

<!-- Sample non-well-formed HTML snippets -->
<div id='MLA1234567' class="container">Some content</div>
<p id="MLA98765432" data-value="123">Another paragraph</p>
<span id=MLA112233445 >Unquoted attribute value</span>
<div id="invalidMLA12345">Too short digits (only 5)</div>
<div class="item" id=" MLA6543210987 ">Valid ID with extra spaces</div>
<div id="MLA123">Way too short (3 digits)</div>

Our goal is to extract IDs that:

  • Are inside an id attribute
  • Start with MLA
  • Followed by 7 or more digits (since "length超过6位" means 7+ digits)

Three Regex Patterns & C# Implementations

Each pattern targets different levels of HTML messiness. All implementations will collect matches into a List<string>.

1. Strict Quoted Attribute Pattern

Best for most cases where id values are wrapped in single or double quotes. This is the most precise and least likely to false-positive.

Regex Pattern:

id=["']MLA(\d{7,})["']

C# Code:

using System.Collections.Generic;
using System.Text.RegularExpressions;

public static List<string> ExtractStrictMlaIds(string htmlText)
{
    var mlaIds = new List<string>();
    const string pattern = @"id=["']MLA(\d{7,})["']";
    
    foreach (Match match in Regex.Matches(htmlText, pattern, RegexOptions.IgnoreCase))
    {
        // Reconstruct the full ID by combining MLA + captured digits
        mlaIds.Add($"MLA{match.Groups[1].Value}");
    }
    
    return mlaIds;
}

2. Unquoted Attribute Support Pattern

Handles non-standard HTML where id values aren't wrapped in quotes (e.g., id=MLA12345678). Uses a backreference to ensure consistent quoting (or no quotes at all).

Regex Pattern:

id\s*=\s*(["']?)MLA(\d{7,})\1

C# Code:

public static List<string> ExtractMlaIdsWithUnquotedSupport(string htmlText)
{
    var mlaIds = new List<string>();
    const string pattern = @"id\s*=\s*(["']?)MLA(\d{7,})\1";
    
    foreach (Match match in Regex.Matches(htmlText, pattern, RegexOptions.IgnoreCase))
    {
        mlaIds.Add($"MLA{match.Groups[2].Value}");
    }
    
    return mlaIds;
}

3. Whitespace-Tolerant Pattern

For extreme cases where id values have extra spaces (e.g., id=" MLA12345678 "). This pattern ignores whitespace inside the attribute value wrapper.

Regex Pattern:

id\s*=\s*(["']?)\s*MLA(\d{7,})\s*\1

C# Code:

public static List<string> ExtractMlaIdsWithWhitespaceTolerance(string htmlText)
{
    var mlaIds = new List<string>();
    const string pattern = @"id\s*=\s*(["']?)\s*MLA(\d{7,})\s*\1";
    
    foreach (Match match in Regex.Matches(htmlText, pattern, RegexOptions.IgnoreCase))
    {
        mlaIds.Add($"MLA{match.Groups[2].Value}");
    }
    
    return mlaIds;
}

Test Usage Example

Here's how you can test these methods with the sample HTML:

public static void Main()
{
    string sampleHtml = @"
<!-- Sample non-well-formed HTML snippets -->
<div id='MLA1234567' class=""container"">Some content</div>
<p id=""MLA98765432"" data-value=""123"">Another paragraph</p>
<span id=MLA112233445 >Unquoted attribute value</span>
<div id=""invalidMLA12345"">Too short digits (only 5)</div>
<div class=""item"" id="" MLA6543210987 "" >Valid ID with extra spaces</div>
<div id=""MLA123"">Way too short (3 digits)</div>
";

    // Test each method
    var strictIds = ExtractStrictMlaIds(sampleHtml);
    var unquotedIds = ExtractMlaIdsWithUnquotedSupport(sampleHtml);
    var whitespaceIds = ExtractMlaIdsWithWhitespaceTolerance(sampleHtml);

    System.Console.WriteLine("Strict Pattern Results:");
    foreach (var id in strictIds) System.Console.WriteLine($"- {id}");

    System.Console.WriteLine("\nUnquoted Support Results:");
    foreach (var id in unquotedIds) System.Console.WriteLine($"- {id}");

    System.Console.WriteLine("\nWhitespace-Tolerant Results:");
    foreach (var id in whitespaceIds) System.Console.WriteLine($"- {id}");
}

Key Notes

  • All patterns use RegexOptions.IgnoreCase to handle cases like ID="MLA1234567" (uppercase ID).
  • If you need to exclude IDs inside HTML comments (<!-- ... -->), you can extend the regex to skip comment blocks, but that adds complexity. Let me know if you need that!
  • Regex isn't perfect for all HTML edge cases, but for non-standard text input, these patterns cover most common scenarios.

内容的提问来源于stack exchange,提问作者Mr John

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:38:26