You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

正则表达式提取字符串中URL,排除结尾句号问题求助

Fixing URL Extraction Regex in PHP

Let's tackle why your current regex isn't working and fix it to handle all your edge cases.

What Was Wrong With Your Original Regex?

Your original pattern had a few key issues:

  1. Too restrictive domain matching: It only allowed letters in domains, missing numbers and hyphens (common in URLs like my-domain123.com).
  2. Captured leading dots: If a URL was attached to preceding text (like sometext.http://google.com), it would include the leading dot in the match.
  3. Word boundary conflicts: The leading \b caused problems with URLs starting with non-word characters (like http://), leading to missed or incorrect matches.
  4. No handling of trailing punctuation: It could accidentally include periods or exclamation marks that were part of the sentence, not the URL.

Revised Solution

Here's an updated regex and PHP code that handles all your scenarios:

for ($i = 0; $i < $resultcount; $i++) {
    // Case-insensitive regex to extract URLs, handling leading dots and trailing punctuation
    $pattern = '%(?:\b\.?)?((?:https?:\/\/)?(?:www\.)?[a-zA-Z0-9-]+(?:\.[a-zA-Z0-9-]+)+(?:\/[^\s,.!?]*)*)(?<![.,])%i';
    $message = (string)$result[$i]['message'];
    preg_match_all($pattern, $message, $matches);
    
    // The cleaned URLs are stored in $matches[1]
    echo "Extracted URLs for post $i:\n";
    print_r($matches[1]);
}

Breakdown of the New Regex

Let's break down what each part does:

  • (?:\b\.?)?: Non-capturing group that handles leading dots from preceding text (like sometext.http://google.com) without including the dot in the extracted URL.
  • (...): The main capture group containing the actual URL:
    • (?:https?:\/\/)?: Matches optional http:// or https:// scheme.
    • (?:www\.)?: Matches optional www. prefix.
    • [a-zA-Z0-9-]+(?:\.[a-zA-Z0-9-]+)+: Matches domain names (supports letters, numbers, hyphens, and multiple subdomain parts like en.wikipedia.org).
    • (?:\/[^\s,.!?]*)*: Matches optional paths, query parameters, or fragments (stops at spaces, commas, or punctuation).
  • (?<![.,]): Negative lookbehind to exclude trailing dots or commas that aren't part of the URL.
  • %i: Case-insensitive modifier (matches HTTP://, HttPs://, etc.).

Testing With Your Example Post

For your sample input:

"This is just a post to test regex for extracting URL. http://google.com , https://www.youtube.com/watch?v=dlw32af https://instagram.com/oscar/ en.wikipedia.org"

The code will extract these URLs:

  • http://google.com
  • https://www.youtube.com/watch?v=dlw32af
  • https://instagram.com/oscar/
  • en.wikipedia.org

It also correctly handles cases like sometext.http://google.com regexDemo., extracting only http://google.com (no leading dot or trailing period).

内容的提问来源于stack exchange,提问作者Mr. Pyramid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:29:49