如何从字典内列表的字符串中移除<start>与<end>标记?
移除字典中所有字符串的和标记
下面两种实用方法可以快速移除字典里所有字符串中的<start>和<end>标记,同时清理掉标记前后的多余空格:
方法1:字符串替换+空白清理
直接用replace()去掉标记,再用strip()清除前后的空白字符,适合标记格式固定的场景:
my_dict = {1:['<start> the woman is slicing onions <end>', '<start> a woman slices a piece of onion with a knife <end>', '<start> a woman is chopping an onion <end>', '<start> a woman is slicing an onion <end>', '<start> a woman is chopping onions <end>', '<start> a woman is slicing onions <end>', '<start> a girl is cutting an onion <end>'], 2: ['<start> a large cat is watching and sniffing a spider <end>', '<start> a cat sniffs a bug <end>', '<start> a cat is sniffing a bug <end>', '<start> a cat is intently watching an insect crawl across the floor <end>', '<start> the cat checked out the bug on the ground <end>', '<start> the cat is watching a bug <end>', '<start> a cat is looking at an insect <end>'], 3:['<start> a man is playing a ukulele <end>', '<start> a man is playing a guitar <end>', '<start> a person is playing a guitar <end>',]} # 处理字典 cleaned_dict = {} for key, sentences in my_dict.items(): cleaned_sentences = [] for s in sentences: # 先移除标记,再清理前后空白 cleaned_s = s.replace('<start>', '').replace('<end>', '').strip() cleaned_sentences.append(cleaned_s) cleaned_dict[key] = cleaned_sentences # 打印验证结果 for k, v in cleaned_dict.items(): print(f"Key {k}:") for sentence in v: print(f" - {sentence}")
方法2:正则表达式(更灵活)
如果标记前后的空格数量不固定,或者需要处理标记的变体,用正则表达式可以一次性匹配并移除标记及周围的空白,扩展性更强:
import re my_dict = {1:['<start> the woman is slicing onions <end>', '<start> a woman slices a piece of onion with a knife <end>', '<start> a woman is chopping an onion <end>', '<start> a woman is slicing an onion <end>', '<start> a woman is chopping onions <end>', '<start> a woman is slicing onions <end>', '<start> a girl is cutting an onion <end>'], 2: ['<start> a large cat is watching and sniffing a spider <end>', '<start> a cat sniffs a bug <end>', '<start> a cat is sniffing a bug <end>', '<start> a cat is intently watching an insect crawl across the floor <end>', '<start> the cat checked out the bug on the ground <end>', '<start> the cat is watching a bug <end>', '<start> a cat is looking at an insect <end>'], 3:['<start> a man is playing a ukulele <end>', '<start> a man is playing a guitar <end>', '<start> a person is playing a guitar <end>',]} # 正则模式:匹配<start>/<end>以及前后任意数量的空白字符 pattern = re.compile(r'\s*<(start|end)>\s*') # 用列表推导式简化处理 cleaned_dict = {key: [pattern.sub('', s) for s in sentences] for key, sentences in my_dict.items()} # 验证结果 for k, v in cleaned_dict.items(): print(f"Key {k}:") for sentence in v: print(f" - {sentence}")
两种方法最终都会得到干净的字符串,你可以根据实际场景选择。
内容的提问来源于stack exchange,提问作者A_B_Y
相关产品推荐
相关产品推荐

