You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python字符串分割保留分隔符及嵌套URL提取技术问询

Hey there! Let's break down these two Python string manipulation tasks for you clearly:

1. Split Strings in Python and Keep Separators in Corresponding Substrings

If you need to split a string while keeping each separator attached to its preceding substring, here are two practical approaches depending on your use case:

Approach 1: Using Regular Expressions (Great for Multiple Separators)

The re.findall() function works perfectly here because we can define a pattern that matches each substring plus its separator (or the final substring without a separator).

import re

def split_with_separators_kept(input_str, separators):
    # Build regex pattern: match substring + separator, or final substring without separator
    regex_pattern = f'(.+?)([{separators}]|$)'
    matched_pairs = re.findall(regex_pattern, input_str)
    # Join each pair and filter out empty strings
    result = [''.join(pair) for pair in matched_pairs if ''.join(pair)]
    return result

# Test it out
test_string = "apple,banana;cherry|date"
print(split_with_separators_kept(test_string, ',;|'))
# Output: ['apple,', 'banana;', 'cherry|', 'date']

Approach 2: Iterative Method (Simple for Single Separator)

If you're only dealing with one separator, a straightforward loop gets the job done without regex:

def split_single_sep_kept(input_str, separator):
    parts = []
    current_part = []
    for char in input_str:
        current_part.append(char)
        if char == separator:
            parts.append(''.join(current_part))
            current_part = []
    # Add the final part if there's anything left
    if current_part:
        parts.append(''.join(current_part))
    return parts

# Test
test_string = "a,b,c,d"
print(split_single_sep_kept(test_string, ','))
# Output: ['a,', 'b,', 'c,', 'd']
2. Extract All Nested URLs from a String

For your specific nested URL string, we can use a regex pattern that targets both standard http/https URLs and double-encoded http URLs (like http%253A%252F%252F which decodes to http:// after two passes).

Here's the code to get exactly the list you need:

import re

# Your input string
url_string = "https://www.link1.net/abc/cik?xai=En8MmT__aF_nQm-F48&sig=Cg0A7_5AE&urlfix=1&;ccurl=https://aax-us.link-two.com/x/c/Qoj_sZnkA%2526adurl%253Dhttp%253A%252F%252Fwww.link-three.mu%252F"

# Regex pattern to match both standard and double-encoded URLs
url_pattern = r'(https?://.*?)(?=https?://|http%253A%252F%252F|$)|(http%253A%252F%252F.*?$)'
matches = re.findall(url_pattern, url_string)

# Combine matches and filter out empty entries
extracted_urls = [''.join(match) for match in matches if ''.join(match)]

print(extracted_urls)
# Output: ['https://www.link1.net/abc/cik?xai=En8MmT__aF_nQm-F48&sig=Cg0A7_5AE&urlfix=1&;ccurl=', 'https://aax-us.link-two.com/x/c/Qoj_sZnkA%2526adurl%253D', 'http%253A%252F%252Fwww.link-three.mu%252F']

How This Works:

  • The first part of the pattern (https?://.*?)(?=https?://|http%253A%252F%252F|$) uses a positive lookahead to match standard URLs until it hits another URL start (either http/https or the double-encoded version) or the end of the string.
  • The second part (http%253A%252F%252F.*?$) captures the double-encoded URL that's at the end of the string.

内容的提问来源于stack exchange,提问作者Adrian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 04:23:49