You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

ECMAScript正则适配解析问题:修正正则并限制循环结果数

Solution for Your ECMAScript Regex & Scraper Issues

Let's break down and fix your problems step by step:

1. Corrected Regex with Required Captures

First, your original regex had a duplicate <td class='s'> which was likely causing it to fail. I've fixed that, plus added captures for seeders, leechers, upload date, and the class="rN" attribute. The regex now aligns with your expected match groups:

var re = /&lt;img src="([\s\S]*?)"[\s\S]*?&lt;td class="nam"&gt;&lt;a href="([\s\S]*?)"[\s\S]*?&gt;([\s\S]*?)&lt;\/a&gt;[\s\S]*?&lt;td class='s'&gt;([\s\S]*?)&lt;\/td&gt;&lt;td class='sl_s'&gt;(\d+)&lt;\/td&gt;&lt;td class='sl_p'&gt;(\d+)&lt;\/td&gt;[\s\S]*?&lt;td class='date'&gt;([\s\S]*?)&lt;\/td&gt;[\s\S]*?(class="r\d+")/g;

Explanation of Match Groups:

  • Match[1]: Image source URL
  • Match[2]: Link href attribute
  • Match[3]: Link text
  • Match[4]: File size (e.g., 6ГБ)
  • Match[5]: Seeders count (e.g., 11)
  • Match[6]: Leechers count (e.g., 0)
  • Match[7]: Upload date/time (e.g., 10.08.2013 в 22:29)
  • Match[8]: Row class (e.g., class="r1")

Key Fixes & Additions:

  • Removed the duplicate &lt;td class='s'&gt; that was breaking the original regex
  • Replaced hardcoded \d with (\d+) to capture seeders/leechers as numeric groups
  • Added a capture for the date field (adjust the class='date' part if your HTML uses a different class name)
  • Added a capture for the class="rN" attribute (adjust the position if this class is on a different element like the parent <tr>)

2. Fixing the Result Limit Loop

Your loop wasn't working because you likely weren't re-executing the regex in each iteration, or not properly incrementing the entry count. Here's the corrected loop logic that limits results to 50:

function scraper_search(html) {
  var re = /&lt;img src="([\s\S]*?)"[\s\S]*?&lt;td class="nam"&gt;&lt;a href="([\s\S]*?)"[\s\S]*?&gt;([\s\S]*?)&lt;\/a&gt;[\s\S]*?&lt;td class='s'&gt;([\s\S]*?)&lt;\/td&gt;&lt;td class='sl_s'&gt;(\d+)&lt;\/td&gt;&lt;td class='sl_p'&gt;(\d+)&lt;\/td&gt;[\s\S]*?&lt;td class='date'&gt;([\s\S]*?)&lt;\/td&gt;[\s\S]*?(class="r\d+")/g;
  var match;
  var entries = [];
  
  // Loop until no more matches or we hit 50 entries
  while ((match = re.exec(html)) !== null && entries.length < 50) {
    entries.push({
      imgSrc: match[1],
      linkHref: match[2],
      linkText: match[3],
      size: match[4],
      seeders: match[5],
      leechers: match[6],
      uploadDate: match[7],
      rowClass: match[8]
    });
  }
  
  return entries;
}

Why This Works:

  • We use re.exec(html) directly in the loop condition to get the next match each iteration
  • We check entries.length < 50 to stop once we've collected 50 results
  • The regex's global flag (g) ensures exec() moves to the next match each time

Notes for Adjustment:

  • If the class="rN" is on the parent <tr> instead of inside the row, adjust the regex to start with &lt;tr (class="r\d+")[\s\S]*? (you'll need to reorder the match groups accordingly)
  • If the date is in a different HTML element (not <td class='date'>), update that part of the regex to match your actual HTML structure

内容的提问来源于stack exchange,提问作者fil brinza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:58:17