如何提取两个指定字符串间的多段目标子串?技术实现求助
Got it, let's fix this substring extraction issue for you. The core problem here is pulling out repeated key-value pairs (name/phone) from a string that looks like a simplified JSON object. Here's a straightforward approach using Python's regex module, which is perfect for this kind of pattern matching:
Step 1: Clean Up the Input String (if needed)
First, notice your input uses Chinese quotation marks (‘/’) instead of standard English ones ('). These will break regex matching, so we'll first convert them to English quotes:
one_string = " ‘name’:’john’,’phone’:’8947467’,’name’:’anthony’,’phone’:’74729’ " # Replace Chinese quotes with English ones cleaned_string = one_string.replace('‘', "'").replace('’', "'")
Step 2: Extract Names and Phones with Regex
We'll use re.findall() to capture all matching values. The regex pattern 'name':'(.*?)' targets the value between 'name':' and the next ' — the .*? is a non-greedy match, which ensures we stop at the first closing quote instead of grabbing everything until the last quote in the string.
Here's the full code:
import re # Original input string one_string = " ‘name’:’john’,’phone’:’8947467’,’name’:’anthony’,’phone’:’74729’ " # Clean up quotation marks cleaned_string = one_string.replace('‘', "'").replace('’', "'") # Extract all names names = re.findall(r"'name':'(.*?)'", cleaned_string) # Extract all phones phones = re.findall(r"'phone':'(.*?)'", cleaned_string) # Format output as you requested formatted_names = ','.join(f"'{name}'" for name in names) formatted_phones = ','.join(f"'{phone}'" for phone in phones) print(f"names = {formatted_names}") print(f"phones = {formatted_phones}")
What This Does:
re.findall()returns a list of all matched values, sonameswill be['john', 'anthony']andphoneswill be['8947467', '74729']- The
join()and f-strings format the lists into the exact string output you want:names = 'john','anthony'andphones = '8947467','74729'
Alternative: Parse as a Dictionary List (More Robust)
If your input string is consistently formatted, you could also convert it into a proper list of dictionaries for more flexibility. Here's how:
import re import ast cleaned_string = one_string.replace('‘', "'").replace('’', "'") # Split into individual key-value pairs and wrap into list of dicts items = re.findall(r"'name':'(.*?)','phone':'(.*?)'", cleaned_string) records = [{'name': name, 'phone': phone} for name, phone in items] names = [record['name'] for record in records] phones = [record['phone'] for record in records]
This approach is better if you need to work with the name-phone pairs together later on.
内容的提问来源于stack exchange,提问作者Johnny

