求助:编写适用于爬虫的Rubular正则表达式匹配download结尾URL
I've put together a regex tailored to your exact needs—one that captures all the URL variants you mentioned, avoids grabbing surrounding HTML, and works with relative, absolute, protocol-less URLs, and those with trailing slashes or query parameters.
Regex for Extracting from HTML href Attributes
If you're scraping URLs directly from <a> tags' href attributes, use this regex to extract only the URL content (not the surrounding HTML):
href=["']\K(?:https?://|www\.|(?!www\.)[a-zA-Z0-9-]+\.[a-zA-Z0-9-.]+/|/)[^\s"']*\/?download(?:\/|\?[^\s"']*)?(?=["'])
Breakdown of the Regex
Let's break down each part so you understand how it works:
href=["']: Matches the start of an HTMLhrefattribute, accounting for both single and double quotes.\K: Resets the match start point, so the precedinghref="orhref='isn't included in your captured result.(?:https?://|www\.|(?!www\.)[a-zA-Z0-9-]+\.[a-zA-Z0-9-.]+/|/): Handles all URL starting scenarios:https?://: Matches HTTP or HTTPS absolute URLswww\.: Matches protocol-less URLs starting withwww.(?!www\.)[a-zA-Z0-9-]+\.[a-zA-Z0-9-.]+/: Matches protocol-less URLs starting with a non-www domain (e.g.,example.com/...)/: Matches relative URLs starting with a slash
[^\s"']*: Matches any path characters until it hits a space, quote, or end of the URL (prevents capturing beyond thehrefvalue)\/?download: Matches the requireddownloadterm, with an optional leading slash (handles paths like/folder/downloador/folder/subfolder/download)(?:\/|\?[^\s"']*)?: Allows optional trailing characters afterdownload:\/: A trailing slash (e.g.,...download/)\?[^\s"']*: A query string (e.g.,...download?cmp=abc)
(?=["']): Ensures the match ends right before the closing quote of thehrefattribute, so no extra HTML is captured.
Regex for Matching Standalone URLs (Not in HTML)
If you need to find these URLs anywhere in text (not just within href attributes), use this variant:
\b(?:https?://|www\.|(?!www\.)[a-zA-Z0-9-]+\.[a-zA-Z0-9-.]+/|/)[^\s<>]*\/?download(?:\/|\?[^\s<>]*)?\b
This uses <> as boundary markers to avoid capturing HTML tag content, while still matching all your target URL types.
Tested Against Your Examples
Both regexes will successfully match all the URL variants you provided:
- Relative URLs:
/product-category/product-name/download,/folder1/download/,/folder1/folder2/download?cmp=abc - Absolute URLs:
https://www.example.com/folder1/download,http://www.example.com/product-category/product-name/download - Protocol-less URLs:
www.example.com/product-category/product-name/download,example.com/folder1/download/
Usage Tips
- Case Insensitivity: If you need to match URLs ending with
Download(capitalized) too, add theiflag (case-insensitive) to your regex settings. - Global Matching: Ensure your tool uses the
gflag to capture all matching URLs on a page, not just the first one. - Validation: While this regex captures all the URL patterns you want, you can post-process the results to filter out invalid URLs (like those that trigger 301s) if needed.
内容的提问来源于stack exchange,提问作者Iam_Amjath

