You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

需求:匹配无分隔符相邻URL的REGEX方案,现有正则失效

Alright, let's figure out how to split those concatenated URLs like www.google.comwww.apple.com into two separate valid links—since your current regex works for standalone URLs but chokes when they're stuck together. Here's a regex-first solution, plus some backup approaches if you need them:

Regex Solution

This regex uses positive lookaheads to detect when one URL ends right before another starts, allowing it to match both URLs correctly even without separators:

\b(?:https?://)?(?:www\.)?[a-zA-Z0-9-]+(?:\.[a-zA-Z0-9-]+)+(?=(?:https?://)?(?:www\.)?[a-zA-Z0-9-]+\.)|\b(?:https?://)?(?:www\.)?[a-zA-Z0-9-]+(?:\.[a-zA-Z0-9-]+)+

How it works:

  • The first half of the regex matches a URL, then uses a positive lookahead (?=(?:https?://)?(?:www\.)?[a-zA-Z0-9-]+\.) to ensure this URL’s end is immediately followed by the start of another valid URL (like www. or a domain prefix with a dot). This catches the first URL in a concatenated pair (e.g., www.google.com in www.google.comwww.apple.com).
  • The second half matches standalone URLs or the remaining URL in a pair (e.g., www.apple.com after the first match is found).

Test case:

For the message Hey checkout these links www.google.comwww.apple.com, this regex will return two matches: www.google.com and www.apple.com.

Alternative Programmatic Approach

If you need more flexibility (like handling URLs with paths or rarer domain suffixes), a small code snippet can complement or replace regex. Here’s a Python example that identifies and splits concatenated URLs:

import re

def split_concatenated_urls(input_text):
    # Base pattern to match valid URLs (adjust suffixes as needed)
    url_pattern = re.compile(r'(?:https?://)?(?:www\.)?[a-zA-Z0-9-]+(?:\.[a-zA-Z0-9-]+)+')
    matches = list(url_pattern.finditer(input_text))
    
    if len(matches) <= 1:
        return [match.group() for match in matches]
    
    result = []
    previous_end = 0
    
    for match in matches:
        current_start, current_end = match.span()
        # Check if this match overlaps with the previous one (sign of concatenated URLs)
        if current_start < previous_end:
            # Split the overlapping part by finding the last valid domain dot
            last_url = result[-1]
            last_dot_pos = last_url.rfind('.')
            if last_dot_pos != -1:
                # Split into the original domain and the new URL's start
                split_part = last_url[last_dot_pos + 1:]
                # Update the last URL and add the new one
                result[-1] = f"{last_url[:last_dot_pos + 1]}{split_part.split('.')[0]}"
                result.append('.'.join(split_part.split('.')[1:]))
        else:
            result.append(match.group())
        
        previous_end = current_end
    
    return result

# Test it out
message = "Hey checkout these links www.google.comwww.apple.com"
print(split_concatenated_urls(message))  # Output: ['www.google.com', 'www.apple.com']

This code first finds all potential URL matches, then checks for overlaps (a sign of concatenated links) and splits them at the logical domain boundary.

内容的提问来源于stack exchange,提问作者Cody E

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:36:36