Python字符串分割保留分隔符及嵌套URL提取技术问询
Hey there! Let's break down these two Python string manipulation tasks for you clearly:
If you need to split a string while keeping each separator attached to its preceding substring, here are two practical approaches depending on your use case:
Approach 1: Using Regular Expressions (Great for Multiple Separators)
The re.findall() function works perfectly here because we can define a pattern that matches each substring plus its separator (or the final substring without a separator).
import re def split_with_separators_kept(input_str, separators): # Build regex pattern: match substring + separator, or final substring without separator regex_pattern = f'(.+?)([{separators}]|$)' matched_pairs = re.findall(regex_pattern, input_str) # Join each pair and filter out empty strings result = [''.join(pair) for pair in matched_pairs if ''.join(pair)] return result # Test it out test_string = "apple,banana;cherry|date" print(split_with_separators_kept(test_string, ',;|')) # Output: ['apple,', 'banana;', 'cherry|', 'date']
Approach 2: Iterative Method (Simple for Single Separator)
If you're only dealing with one separator, a straightforward loop gets the job done without regex:
def split_single_sep_kept(input_str, separator): parts = [] current_part = [] for char in input_str: current_part.append(char) if char == separator: parts.append(''.join(current_part)) current_part = [] # Add the final part if there's anything left if current_part: parts.append(''.join(current_part)) return parts # Test test_string = "a,b,c,d" print(split_single_sep_kept(test_string, ',')) # Output: ['a,', 'b,', 'c,', 'd']
For your specific nested URL string, we can use a regex pattern that targets both standard http/https URLs and double-encoded http URLs (like http%253A%252F%252F which decodes to http:// after two passes).
Here's the code to get exactly the list you need:
import re # Your input string url_string = "https://www.link1.net/abc/cik?xai=En8MmT__aF_nQm-F48&sig=Cg0A7_5AE&urlfix=1&;ccurl=https://aax-us.link-two.com/x/c/Qoj_sZnkA%2526adurl%253Dhttp%253A%252F%252Fwww.link-three.mu%252F" # Regex pattern to match both standard and double-encoded URLs url_pattern = r'(https?://.*?)(?=https?://|http%253A%252F%252F|$)|(http%253A%252F%252F.*?$)' matches = re.findall(url_pattern, url_string) # Combine matches and filter out empty entries extracted_urls = [''.join(match) for match in matches if ''.join(match)] print(extracted_urls) # Output: ['https://www.link1.net/abc/cik?xai=En8MmT__aF_nQm-F48&sig=Cg0A7_5AE&urlfix=1&;ccurl=', 'https://aax-us.link-two.com/x/c/Qoj_sZnkA%2526adurl%253D', 'http%253A%252F%252Fwww.link-three.mu%252F']
How This Works:
- The first part of the pattern
(https?://.*?)(?=https?://|http%253A%252F%252F|$)uses a positive lookahead to match standard URLs until it hits another URL start (eitherhttp/httpsor the double-encoded version) or the end of the string. - The second part
(http%253A%252F%252F.*?$)captures the double-encoded URL that's at the end of the string.
内容的提问来源于stack exchange,提问作者Adrian

