ECMAScript正则适配解析问题:修正正则并限制循环结果数
Solution for Your ECMAScript Regex & Scraper Issues
Let's break down and fix your problems step by step:
1. Corrected Regex with Required Captures
First, your original regex had a duplicate <td class='s'> which was likely causing it to fail. I've fixed that, plus added captures for seeders, leechers, upload date, and the class="rN" attribute. The regex now aligns with your expected match groups:
var re = /<img src="([\s\S]*?)"[\s\S]*?<td class="nam"><a href="([\s\S]*?)"[\s\S]*?>([\s\S]*?)<\/a>[\s\S]*?<td class='s'>([\s\S]*?)<\/td><td class='sl_s'>(\d+)<\/td><td class='sl_p'>(\d+)<\/td>[\s\S]*?<td class='date'>([\s\S]*?)<\/td>[\s\S]*?(class="r\d+")/g;
Explanation of Match Groups:
- Match[1]: Image source URL
- Match[2]: Link href attribute
- Match[3]: Link text
- Match[4]: File size (e.g.,
6ГБ) - Match[5]: Seeders count (e.g.,
11) - Match[6]: Leechers count (e.g.,
0) - Match[7]: Upload date/time (e.g.,
10.08.2013 в 22:29) - Match[8]: Row class (e.g.,
class="r1")
Key Fixes & Additions:
- Removed the duplicate
<td class='s'>that was breaking the original regex - Replaced hardcoded
\dwith(\d+)to capture seeders/leechers as numeric groups - Added a capture for the date field (adjust the
class='date'part if your HTML uses a different class name) - Added a capture for the
class="rN"attribute (adjust the position if this class is on a different element like the parent<tr>)
2. Fixing the Result Limit Loop
Your loop wasn't working because you likely weren't re-executing the regex in each iteration, or not properly incrementing the entry count. Here's the corrected loop logic that limits results to 50:
function scraper_search(html) { var re = /<img src="([\s\S]*?)"[\s\S]*?<td class="nam"><a href="([\s\S]*?)"[\s\S]*?>([\s\S]*?)<\/a>[\s\S]*?<td class='s'>([\s\S]*?)<\/td><td class='sl_s'>(\d+)<\/td><td class='sl_p'>(\d+)<\/td>[\s\S]*?<td class='date'>([\s\S]*?)<\/td>[\s\S]*?(class="r\d+")/g; var match; var entries = []; // Loop until no more matches or we hit 50 entries while ((match = re.exec(html)) !== null && entries.length < 50) { entries.push({ imgSrc: match[1], linkHref: match[2], linkText: match[3], size: match[4], seeders: match[5], leechers: match[6], uploadDate: match[7], rowClass: match[8] }); } return entries; }
Why This Works:
- We use
re.exec(html)directly in the loop condition to get the next match each iteration - We check
entries.length < 50to stop once we've collected 50 results - The regex's global flag (
g) ensuresexec()moves to the next match each time
Notes for Adjustment:
- If the
class="rN"is on the parent<tr>instead of inside the row, adjust the regex to start with<tr (class="r\d+")[\s\S]*?(you'll need to reorder the match groups accordingly) - If the date is in a different HTML element (not
<td class='date'>), update that part of the regex to match your actual HTML structure
内容的提问来源于stack exchange,提问作者fil brinza
相关产品推荐
相关产品推荐

