You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:编写适用于爬虫的Rubular正则表达式匹配download结尾URL

Solution for Matching URLs Ending with "download"

I've put together a regex tailored to your exact needs—one that captures all the URL variants you mentioned, avoids grabbing surrounding HTML, and works with relative, absolute, protocol-less URLs, and those with trailing slashes or query parameters.

Regex for Extracting from HTML href Attributes

If you're scraping URLs directly from <a> tags' href attributes, use this regex to extract only the URL content (not the surrounding HTML):

href=["']\K(?:https?://|www\.|(?!www\.)[a-zA-Z0-9-]+\.[a-zA-Z0-9-.]+/|/)[^\s"']*\/?download(?:\/|\?[^\s"']*)?(?=["'])

Breakdown of the Regex

Let's break down each part so you understand how it works:

  • href=["']: Matches the start of an HTML href attribute, accounting for both single and double quotes.
  • \K: Resets the match start point, so the preceding href=" or href=' isn't included in your captured result.
  • (?:https?://|www\.|(?!www\.)[a-zA-Z0-9-]+\.[a-zA-Z0-9-.]+/|/): Handles all URL starting scenarios:
    • https?://: Matches HTTP or HTTPS absolute URLs
    • www\.: Matches protocol-less URLs starting with www.
    • (?!www\.)[a-zA-Z0-9-]+\.[a-zA-Z0-9-.]+/: Matches protocol-less URLs starting with a non-www domain (e.g., example.com/...)
    • /: Matches relative URLs starting with a slash
  • [^\s"']*: Matches any path characters until it hits a space, quote, or end of the URL (prevents capturing beyond the href value)
  • \/?download: Matches the required download term, with an optional leading slash (handles paths like /folder/download or /folder/subfolder/download)
  • (?:\/|\?[^\s"']*)?: Allows optional trailing characters after download:
    • \/: A trailing slash (e.g., ...download/)
    • \?[^\s"']*: A query string (e.g., ...download?cmp=abc)
  • (?=["']): Ensures the match ends right before the closing quote of the href attribute, so no extra HTML is captured.

Regex for Matching Standalone URLs (Not in HTML)

If you need to find these URLs anywhere in text (not just within href attributes), use this variant:

\b(?:https?://|www\.|(?!www\.)[a-zA-Z0-9-]+\.[a-zA-Z0-9-.]+/|/)[^\s<>]*\/?download(?:\/|\?[^\s<>]*)?\b

This uses <> as boundary markers to avoid capturing HTML tag content, while still matching all your target URL types.

Tested Against Your Examples

Both regexes will successfully match all the URL variants you provided:

  • Relative URLs: /product-category/product-name/download, /folder1/download/, /folder1/folder2/download?cmp=abc
  • Absolute URLs: https://www.example.com/folder1/download, http://www.example.com/product-category/product-name/download
  • Protocol-less URLs: www.example.com/product-category/product-name/download, example.com/folder1/download/

Usage Tips

  • Case Insensitivity: If you need to match URLs ending with Download (capitalized) too, add the i flag (case-insensitive) to your regex settings.
  • Global Matching: Ensure your tool uses the g flag to capture all matching URLs on a page, not just the first one.
  • Validation: While this regex captures all the URL patterns you want, you can post-process the results to filter out invalid URLs (like those that trigger 301s) if needed.

内容的提问来源于stack exchange,提问作者Iam_Amjath

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:02:58