基于字符串模板与差异字符串列表生成字典的最优实现方案咨询
Hey there! Great question—this is a super common task when working with structured string templates, and we can build a flexible, reusable solution that adapts to any template using the {variable} format you're working with. Let's walk through how to do this properly.
Step 1: Convert Your Template to a Regex Pattern
First, we need to turn your human-readable template into a regular expression that can capture the variable parts. The key here is to:
- Escape any special regex characters in the template (like
?,-, or.which have special meaning in regex) - Replace each
{variable}with a named capture group so we can easily map matches to their variable names
Here's a function to handle that:
import re def template_to_regex(template): # Escape all regex-special characters in the template first escaped_template = re.escape(template) # Replace escaped `{var}` patterns with named regex capture groups regex_pattern = re.sub(r'\\\{(.*?)\\\}', r'(?P<\1>.*?)', escaped_template) # Compile the pattern for faster matching return re.compile(regex_pattern)
Step 2: Extract and Transform Matched Data
Next, we'll write a function to match each input string against the regex, pull out the captured variables, and apply any necessary type conversions (like turning your comma-separated list_of_data into an integer list).
def extract_data(template_regex, input_string, type_converters=None): # Try to match the input string against our template regex match = template_regex.match(input_string) if not match: return None # You could also raise an error here if mismatched strings are a problem # Get a dictionary of variable names to their raw string values result = match.groupdict() # Apply custom type converters if provided if type_converters: for key, converter in type_converters.items(): if key in result: result[key] = converter(result[key]) return result
Step 3: Put It All Together for Your Example
Now let's use these functions with your specific template and input strings. We'll define a converter for list_of_data to turn the comma-separated string into a list of integers:
# Your original template and input strings template = "Hi {name}, how are you? Are you living in {location} currently? Can you confirm if following data is correct - {list_of_data}" list_of_strings = [ "Hi John, how are you? Are you living in California currently? Can you confirm if following data is correct - 123, 456, 345", "Hi Steve, how are you? Are you living in New York currently? Can you confirm if following data is correct - 6542" ] # Define a converter for list_of_data def convert_list_data(raw_string): # Split by commas, strip whitespace, and convert each item to int return [int(item.strip()) for item in raw_string.split(',')] # Map variable names to their converters type_converters = { "list_of_data": convert_list_data } # Create our regex pattern from the template template_regex = template_to_regex(template) # Process each input string output = [] for string in list_of_strings: extracted = extract_data(template_regex, string, type_converters) if extracted: output.append(extracted) print(output)
Output
Running this code will give you exactly the expected result:
[ {"name": "John", "location": "California", "list_of_data": [123, 456, 345]}, {"name": "Steve", "location": "New York", "list_of_data": [6542]} ]
Why This Is a Generic Solution
This approach works for any template using {variable} syntax:
- You don't have to write custom string slicing or splitting logic for each new template
- The
type_convertersdictionary lets you handle different data types (dates, floats, booleans) easily - It handles all regex-special characters in your template automatically
Notes
- If an input string doesn't match the template, the function returns
None—you can adjust this to raise an exception or log a warning if needed - The regex uses
.*?(non-greedy matching) to ensure it doesn't accidentally capture extra characters between variables
内容的提问来源于stack exchange,提问作者NeelamS

