You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于字符串模板与差异字符串列表生成字典的最优实现方案咨询

Generic Solution to Extract Variables from String Templates

Hey there! Great question—this is a super common task when working with structured string templates, and we can build a flexible, reusable solution that adapts to any template using the {variable} format you're working with. Let's walk through how to do this properly.

Step 1: Convert Your Template to a Regex Pattern

First, we need to turn your human-readable template into a regular expression that can capture the variable parts. The key here is to:

  • Escape any special regex characters in the template (like ?, -, or . which have special meaning in regex)
  • Replace each {variable} with a named capture group so we can easily map matches to their variable names

Here's a function to handle that:

import re

def template_to_regex(template):
    # Escape all regex-special characters in the template first
    escaped_template = re.escape(template)
    # Replace escaped `{var}` patterns with named regex capture groups
    regex_pattern = re.sub(r'\\\{(.*?)\\\}', r'(?P<\1>.*?)', escaped_template)
    # Compile the pattern for faster matching
    return re.compile(regex_pattern)

Step 2: Extract and Transform Matched Data

Next, we'll write a function to match each input string against the regex, pull out the captured variables, and apply any necessary type conversions (like turning your comma-separated list_of_data into an integer list).

def extract_data(template_regex, input_string, type_converters=None):
    # Try to match the input string against our template regex
    match = template_regex.match(input_string)
    if not match:
        return None  # You could also raise an error here if mismatched strings are a problem
    
    # Get a dictionary of variable names to their raw string values
    result = match.groupdict()
    
    # Apply custom type converters if provided
    if type_converters:
        for key, converter in type_converters.items():
            if key in result:
                result[key] = converter(result[key])
    
    return result

Step 3: Put It All Together for Your Example

Now let's use these functions with your specific template and input strings. We'll define a converter for list_of_data to turn the comma-separated string into a list of integers:

# Your original template and input strings
template = "Hi {name}, how are you? Are you living in {location} currently? Can you confirm if following data is correct - {list_of_data}"
list_of_strings = [ 
    "Hi John, how are you? Are you living in California currently? Can you confirm if following data is correct - 123, 456, 345",
    "Hi Steve, how are you? Are you living in New York currently? Can you confirm if following data is correct - 6542"
]

# Define a converter for list_of_data
def convert_list_data(raw_string):
    # Split by commas, strip whitespace, and convert each item to int
    return [int(item.strip()) for item in raw_string.split(',')]

# Map variable names to their converters
type_converters = {
    "list_of_data": convert_list_data
}

# Create our regex pattern from the template
template_regex = template_to_regex(template)

# Process each input string
output = []
for string in list_of_strings:
    extracted = extract_data(template_regex, string, type_converters)
    if extracted:
        output.append(extracted)

print(output)

Output

Running this code will give you exactly the expected result:

[
    {"name": "John", "location": "California", "list_of_data": [123, 456, 345]},
    {"name": "Steve", "location": "New York", "list_of_data": [6542]}
]

Why This Is a Generic Solution

This approach works for any template using {variable} syntax:

  • You don't have to write custom string slicing or splitting logic for each new template
  • The type_converters dictionary lets you handle different data types (dates, floats, booleans) easily
  • It handles all regex-special characters in your template automatically

Notes

  • If an input string doesn't match the template, the function returns None—you can adjust this to raise an exception or log a warning if needed
  • The regex uses .*? (non-greedy matching) to ensure it doesn't accidentally capture extra characters between variables

内容的提问来源于stack exchange,提问作者NeelamS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 18:28:13