You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从字典内列表的字符串中移除<start>与<end>标记?

移除字典中所有字符串的和标记

下面两种实用方法可以快速移除字典里所有字符串中的<start>和<end>标记,同时清理掉标记前后的多余空格:

方法1:字符串替换+空白清理

直接用replace()去掉标记,再用strip()清除前后的空白字符,适合标记格式固定的场景:

my_dict = {1:['<start> the woman is slicing onions <end>',
       '<start> a woman slices a piece of onion with a knife <end>',
       '<start> a woman is chopping an onion <end>',
       '<start> a woman is slicing an onion <end>',
       '<start> a woman is chopping onions <end>',
       '<start> a woman is slicing onions <end>',
       '<start> a girl is cutting an onion <end>'],
      2: ['<start> a large cat is watching and sniffing a spider <end>',
            '<start> a cat sniffs a bug <end>',
            '<start> a cat is sniffing a bug <end>',
            '<start> a cat is intently watching an insect crawl across the floor <end>',
            '<start> the cat checked out the bug on the ground <end>',
            '<start> the cat is watching a bug <end>',
            '<start> a cat is looking at an insect <end>'],
      3:['<start> a man is playing a ukulele <end>',
         '<start> a man is playing a guitar <end>',
         '<start> a person is playing a guitar <end>',]}

# 处理字典
cleaned_dict = {}
for key, sentences in my_dict.items():
    cleaned_sentences = []
    for s in sentences:
        # 先移除标记,再清理前后空白
        cleaned_s = s.replace('<start>', '').replace('<end>', '').strip()
        cleaned_sentences.append(cleaned_s)
    cleaned_dict[key] = cleaned_sentences

# 打印验证结果
for k, v in cleaned_dict.items():
    print(f"Key {k}:")
    for sentence in v:
        print(f"  - {sentence}")

方法2:正则表达式(更灵活)

如果标记前后的空格数量不固定,或者需要处理标记的变体,用正则表达式可以一次性匹配并移除标记及周围的空白,扩展性更强:

import re

my_dict = {1:['<start> the woman is slicing onions <end>',
       '<start> a woman slices a piece of onion with a knife <end>',
       '<start> a woman is chopping an onion <end>',
       '<start> a woman is slicing an onion <end>',
       '<start> a woman is chopping onions <end>',
       '<start> a woman is slicing onions <end>',
       '<start> a girl is cutting an onion <end>'],
      2: ['<start> a large cat is watching and sniffing a spider <end>',
            '<start> a cat sniffs a bug <end>',
            '<start> a cat is sniffing a bug <end>',
            '<start> a cat is intently watching an insect crawl across the floor <end>',
            '<start> the cat checked out the bug on the ground <end>',
            '<start> the cat is watching a bug <end>',
            '<start> a cat is looking at an insect <end>'],
      3:['<start> a man is playing a ukulele <end>',
         '<start> a man is playing a guitar <end>',
         '<start> a person is playing a guitar <end>',]}

# 正则模式:匹配<start>/<end>以及前后任意数量的空白字符
pattern = re.compile(r'\s*<(start|end)>\s*')

# 用列表推导式简化处理
cleaned_dict = {key: [pattern.sub('', s) for s in sentences] 
                for key, sentences in my_dict.items()}

# 验证结果
for k, v in cleaned_dict.items():
    print(f"Key {k}:")
    for sentence in v:
        print(f"  - {sentence}")

两种方法最终都会得到干净的字符串,你可以根据实际场景选择。

内容的提问来源于stack exchange,提问作者A_B_Y

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 13:25:18