You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式MatchCollection失效问题求助(HTML提取场景)

Hey there! No worries about your English at all—let's dig into why your Extract method works with test data but fails on real HTML. Here are the most likely issues and fixes:

1. You're matching HTML entities instead of actual tags

Your test data uses &lt; and &gt; (the HTML entity versions of < and >), so your current Extract call looks for those entities. But real HTML from websites uses the actual <li> and </li> tags, not their entity equivalents. That's the biggest reason your regex isn't finding matches!

Fix: Update your Extract call to use real tag characters instead of entities:

// Before:
// str = Extract(www.text, "&lt;li class=\"a-class\"&gt;", "&lt;/li&gt;");

// After:
str = Extract(www.text, "<li class=\"a-class\">", "</li>");

2. Real HTML has messy formatting (spaces, extra attributes)

Test data is perfectly clean, but real HTML often has extra spaces, line breaks, or additional attributes in tags (like <li class="a-class" id="item-1">). Your current regex won't account for that.

Fix: Make your regex more flexible to handle these variations. Update the Extract method's regex pattern:

// Use verbatim string (@) to avoid escaping headaches, add flexibility for spaces/quotes
Regex regex = new Regex(@"(?<=<li\s+class=[""]a-class[""]\s*>)(.*?)(?=</li>)", 
                        RegexOptions.Singleline | RegexOptions.IgnoreCase);
  • \s+ matches one or more spaces (including newlines/tabs)
  • [""] matches double quotes (use ["'] if the site might use single quotes for classes)
  • RegexOptions.Singleline lets . match line breaks (in case <li> content spans lines)
  • RegexOptions.IgnoreCase handles accidental capitalization in tags (optional, but safe)

3. HTML decoding timing might be off

You tried HttpUtility.HtmlDecode, but if you decode after extracting, you're working with encoded content. Decode the full HTML string first before running your extraction:

string decodedHtml = HttpUtility.HtmlDecode(www.text);
str = Extract(decodedHtml, "<li class=\"a-class\">", "</li>");

Bonus: Stop using regex for HTML (it's unreliable!)

Regex is great for simple patterns, but HTML is a structured language with nested tags, comments, and weird edge cases that regex can't handle consistently. For a more robust solution, use an HTML parsing library like HtmlAgilityPack (it's free and works great with Unity).

Example with HtmlAgilityPack:

  1. Install the HtmlAgilityPack NuGet package in your project.
  2. Rewrite your GetHtml coroutine like this:
IEnumerator GetHtml() {
    WWW www = new WWW("https://www.random-site.com/" + search.text); 
    yield return www; 

    if (www.text.Contains("No search results were found for")) {
        Debug.LogError("Aucun résultat pour \"" + search.text + "\"");
        StopCoroutine(AllMusic());
        yield break;
    }

    string decodedHtml = HttpUtility.HtmlDecode(www.text);
    HtmlDocument doc = new HtmlDocument();
    doc.LoadHtml(decodedHtml);

    // Find all <li> tags with class "a-class"
    var liNodes = doc.DocumentNode.SelectNodes("//li[@class='a-class']");
    List<string> results = new List<string>();

    if (liNodes != null) {
        foreach (var li in liNodes) {
            // Get the text inside the <li>, trim extra spaces/newlines
            results.Add(li.InnerText.Trim());
        }
    }

    File.WriteAllLines(@"azaaac.txt", results.ToArray());
}

This will handle any formatting quirks, nested tags, or attribute variations automatically.

Give these fixes a try—start with changing the entity tags to real ones first, since that's the most likely culprit!

内容的提问来源于stack exchange,提问作者Aiixu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:57:11