正则表达式MatchCollection失效问题求助(HTML提取场景)
Hey there! No worries about your English at all—let's dig into why your Extract method works with test data but fails on real HTML. Here are the most likely issues and fixes:
1. You're matching HTML entities instead of actual tags
Your test data uses < and > (the HTML entity versions of < and >), so your current Extract call looks for those entities. But real HTML from websites uses the actual <li> and </li> tags, not their entity equivalents. That's the biggest reason your regex isn't finding matches!
Fix: Update your Extract call to use real tag characters instead of entities:
// Before: // str = Extract(www.text, "<li class=\"a-class\">", "</li>"); // After: str = Extract(www.text, "<li class=\"a-class\">", "</li>");
2. Real HTML has messy formatting (spaces, extra attributes)
Test data is perfectly clean, but real HTML often has extra spaces, line breaks, or additional attributes in tags (like <li class="a-class" id="item-1">). Your current regex won't account for that.
Fix: Make your regex more flexible to handle these variations. Update the Extract method's regex pattern:
// Use verbatim string (@) to avoid escaping headaches, add flexibility for spaces/quotes Regex regex = new Regex(@"(?<=<li\s+class=[""]a-class[""]\s*>)(.*?)(?=</li>)", RegexOptions.Singleline | RegexOptions.IgnoreCase);
\s+matches one or more spaces (including newlines/tabs)[""]matches double quotes (use["']if the site might use single quotes for classes)RegexOptions.Singlelinelets.match line breaks (in case<li>content spans lines)RegexOptions.IgnoreCasehandles accidental capitalization in tags (optional, but safe)
3. HTML decoding timing might be off
You tried HttpUtility.HtmlDecode, but if you decode after extracting, you're working with encoded content. Decode the full HTML string first before running your extraction:
string decodedHtml = HttpUtility.HtmlDecode(www.text); str = Extract(decodedHtml, "<li class=\"a-class\">", "</li>");
Bonus: Stop using regex for HTML (it's unreliable!)
Regex is great for simple patterns, but HTML is a structured language with nested tags, comments, and weird edge cases that regex can't handle consistently. For a more robust solution, use an HTML parsing library like HtmlAgilityPack (it's free and works great with Unity).
Example with HtmlAgilityPack:
- Install the HtmlAgilityPack NuGet package in your project.
- Rewrite your
GetHtmlcoroutine like this:
IEnumerator GetHtml() { WWW www = new WWW("https://www.random-site.com/" + search.text); yield return www; if (www.text.Contains("No search results were found for")) { Debug.LogError("Aucun résultat pour \"" + search.text + "\""); StopCoroutine(AllMusic()); yield break; } string decodedHtml = HttpUtility.HtmlDecode(www.text); HtmlDocument doc = new HtmlDocument(); doc.LoadHtml(decodedHtml); // Find all <li> tags with class "a-class" var liNodes = doc.DocumentNode.SelectNodes("//li[@class='a-class']"); List<string> results = new List<string>(); if (liNodes != null) { foreach (var li in liNodes) { // Get the text inside the <li>, trim extra spaces/newlines results.Add(li.InnerText.Trim()); } } File.WriteAllLines(@"azaaac.txt", results.ToArray()); }
This will handle any formatting quirks, nested tags, or attribute variations automatically.
Give these fixes a try—start with changing the entity tags to real ones first, since that's the most likely culprit!
内容的提问来源于stack exchange,提问作者Aiixu

