需求:匹配无分隔符相邻URL的REGEX方案,现有正则失效
Alright, let's figure out how to split those concatenated URLs like www.google.comwww.apple.com into two separate valid links—since your current regex works for standalone URLs but chokes when they're stuck together. Here's a regex-first solution, plus some backup approaches if you need them:
This regex uses positive lookaheads to detect when one URL ends right before another starts, allowing it to match both URLs correctly even without separators:
\b(?:https?://)?(?:www\.)?[a-zA-Z0-9-]+(?:\.[a-zA-Z0-9-]+)+(?=(?:https?://)?(?:www\.)?[a-zA-Z0-9-]+\.)|\b(?:https?://)?(?:www\.)?[a-zA-Z0-9-]+(?:\.[a-zA-Z0-9-]+)+
How it works:
- The first half of the regex matches a URL, then uses a positive lookahead
(?=(?:https?://)?(?:www\.)?[a-zA-Z0-9-]+\.)to ensure this URL’s end is immediately followed by the start of another valid URL (likewww.or a domain prefix with a dot). This catches the first URL in a concatenated pair (e.g.,www.google.cominwww.google.comwww.apple.com). - The second half matches standalone URLs or the remaining URL in a pair (e.g.,
www.apple.comafter the first match is found).
Test case:
For the message Hey checkout these links www.google.comwww.apple.com, this regex will return two matches: www.google.com and www.apple.com.
If you need more flexibility (like handling URLs with paths or rarer domain suffixes), a small code snippet can complement or replace regex. Here’s a Python example that identifies and splits concatenated URLs:
import re def split_concatenated_urls(input_text): # Base pattern to match valid URLs (adjust suffixes as needed) url_pattern = re.compile(r'(?:https?://)?(?:www\.)?[a-zA-Z0-9-]+(?:\.[a-zA-Z0-9-]+)+') matches = list(url_pattern.finditer(input_text)) if len(matches) <= 1: return [match.group() for match in matches] result = [] previous_end = 0 for match in matches: current_start, current_end = match.span() # Check if this match overlaps with the previous one (sign of concatenated URLs) if current_start < previous_end: # Split the overlapping part by finding the last valid domain dot last_url = result[-1] last_dot_pos = last_url.rfind('.') if last_dot_pos != -1: # Split into the original domain and the new URL's start split_part = last_url[last_dot_pos + 1:] # Update the last URL and add the new one result[-1] = f"{last_url[:last_dot_pos + 1]}{split_part.split('.')[0]}" result.append('.'.join(split_part.split('.')[1:])) else: result.append(match.group()) previous_end = current_end return result # Test it out message = "Hey checkout these links www.google.comwww.apple.com" print(split_concatenated_urls(message)) # Output: ['www.google.com', 'www.apple.com']
This code first finds all potential URL matches, then checks for overlaps (a sign of concatenated links) and splits them at the logical domain boundary.
内容的提问来源于stack exchange,提问作者Cody E

