如何优雅实现基于text标签的字符串提取及变量替换?
实现方案
要优雅完成这个需求,可以分三步处理:提取变量映射、提取text字段内容、替换变量占位符,以下是具体实现思路和代码:
步骤1:提取VAR变量映射
先用正则匹配所有<VAR_*>标签及其内容,构建变量名到文本的字典,同时处理标签内的换行和首尾空白,保证提取的文本干净。
步骤2:提取text字段内容
先定位<interaction>标签内的JSON内容,解析JSON后逐层提取speech数组中的text值;如果JSON解析失败(比如格式转义问题),则降级用正则匹配所有"text": "..."格式的内容,增强鲁棒性。
步骤3:替换变量占位符
对每个提取到的text内容,用正则匹配{{{VAR_(\d+)}}}格式的占位符,替换成对应的变量文本;如果找不到对应变量,保留原占位符(可根据需求调整处理逻辑)。
Python 实现代码
import re import json def extract_and_replace_text(content): # 提取VAR变量映射 var_pattern = re.compile(r'<VAR_(\d+)>\s*(.*?)\s*</VAR_\1>', re.DOTALL) var_map = {} for match in var_pattern.finditer(content): var_name = f"VAR_{match.group(1)}" var_text = match.group(2).strip() var_map[var_name] = var_text # 提取<interaction>内的内容 interaction_pattern = re.compile(r'<interaction.*?>\s*(.*?)\s*</interaction>', re.DOTALL) interaction_match = interaction_pattern.search(content) if not interaction_match: return None texts = [] try: # 处理转义引号并解析JSON json_content = interaction_match.group(1).replace('"', '"') data = json.loads(json_content) speech_list = data.get("BaseClassConfig QuestionInteractionInterface", {})\ .get("QuestionSpeakingInteraction", {})\ .get("SpeakingInteractionConfig", {})\ .get("speech", []) texts = [item.get("text", "") for item in speech_list] except json.JSONDecodeError: # JSON解析失败时用正则匹配text字段 text_pattern = re.compile(r'"text":\s*"(.*?)"', re.DOTALL) raw_texts = text_pattern.findall(interaction_match.group(1).replace('"', '"')) texts = [t.strip() for t in raw_texts] if not texts: return None # 替换变量占位符 def replace_var(match): var_name = match.group(1) return var_map.get(var_name, match.group(0)) result = [re.sub(r'{{{(VAR_\d+)}}}', replace_var, text) for text in texts] return result # 测试示例内容 sample_content = """ <VAR_0> How old are you? </VAR_0> <VAR_2> What is your age? </VAR_2> (set:$ageask to 'true') (set:$storybeat_01 to 1) (set:$inputAge to -1) <audio stopall> <audio src="sound.wav" autoplay loop volume="100%"> <onSuccess($user_age==0,$inputAge:=age)>[[age selector]] <onSuccess(5>age OR 100<age,$inputAge:=age)>[[impossible age]] <onSuccess(90<age,$inputAge:=age)>[[very old]] <interaction type="AgeNumberQuestionInteraction"> { "BaseClassConfig QuestionInteractionInterface": { "QuestionInteractionInterfaceConfig": { "maxQuestionRepetitionCount": 1 }, "QuestionSpeakingInteraction": { "SpeakingInteractionConfig": { "speech": [ { "emotion": "HAPPY", "text": "{{{VAR_2}}}" }, { "emotion": "HAPPY", "text": "Tell me. How old are you?" }, { "emotion": "HAPPY", "text": "{{{VAR_0}}}" } ] } } } } </interaction> """ # 执行测试 print(extract_and_replace_text(sample_content)) # 输出: ['What is your age?', 'Tell me. How old are you?', 'How old are you?']
关键说明
- 变量提取使用
re.DOTALL模式支持跨换行匹配,strip()清理文本首尾冗余空白。 - 优先用JSON解析保证文本提取的准确性,正则作为降级方案适配格式异常情况。
- 占位符替换用回调函数实现,逻辑清晰,便于后续扩展(比如增加变量不存在时的默认值处理)。
- 无text字段时返回
None,完全符合需求定义。
内容的提问来源于stack exchange,提问作者BlackHawk
相关产品推荐
相关产品推荐

