正则表达式提取字符串中URL,排除结尾句号问题求助
Fixing URL Extraction Regex in PHP
Let's tackle why your current regex isn't working and fix it to handle all your edge cases.
What Was Wrong With Your Original Regex?
Your original pattern had a few key issues:
- Too restrictive domain matching: It only allowed letters in domains, missing numbers and hyphens (common in URLs like
my-domain123.com). - Captured leading dots: If a URL was attached to preceding text (like
sometext.http://google.com), it would include the leading dot in the match. - Word boundary conflicts: The leading
\bcaused problems with URLs starting with non-word characters (likehttp://), leading to missed or incorrect matches. - No handling of trailing punctuation: It could accidentally include periods or exclamation marks that were part of the sentence, not the URL.
Revised Solution
Here's an updated regex and PHP code that handles all your scenarios:
for ($i = 0; $i < $resultcount; $i++) { // Case-insensitive regex to extract URLs, handling leading dots and trailing punctuation $pattern = '%(?:\b\.?)?((?:https?:\/\/)?(?:www\.)?[a-zA-Z0-9-]+(?:\.[a-zA-Z0-9-]+)+(?:\/[^\s,.!?]*)*)(?<![.,])%i'; $message = (string)$result[$i]['message']; preg_match_all($pattern, $message, $matches); // The cleaned URLs are stored in $matches[1] echo "Extracted URLs for post $i:\n"; print_r($matches[1]); }
Breakdown of the New Regex
Let's break down what each part does:
(?:\b\.?)?: Non-capturing group that handles leading dots from preceding text (likesometext.http://google.com) without including the dot in the extracted URL.(...): The main capture group containing the actual URL:(?:https?:\/\/)?: Matches optionalhttp://orhttps://scheme.(?:www\.)?: Matches optionalwww.prefix.[a-zA-Z0-9-]+(?:\.[a-zA-Z0-9-]+)+: Matches domain names (supports letters, numbers, hyphens, and multiple subdomain parts likeen.wikipedia.org).(?:\/[^\s,.!?]*)*: Matches optional paths, query parameters, or fragments (stops at spaces, commas, or punctuation).
(?<![.,]): Negative lookbehind to exclude trailing dots or commas that aren't part of the URL.%i: Case-insensitive modifier (matchesHTTP://,HttPs://, etc.).
Testing With Your Example Post
For your sample input:
"This is just a post to test regex for extracting URL. http://google.com , https://www.youtube.com/watch?v=dlw32af https://instagram.com/oscar/ en.wikipedia.org"
The code will extract these URLs:
http://google.comhttps://www.youtube.com/watch?v=dlw32afhttps://instagram.com/oscar/en.wikipedia.org
It also correctly handles cases like sometext.http://google.com regexDemo., extracting only http://google.com (no leading dot or trailing period).
内容的提问来源于stack exchange,提问作者Mr. Pyramid
相关产品推荐
相关产品推荐

