如何提取列表中被#符号包裹的成对内容集合?
解决思路
核心用状态标记区分是否处于两个#包裹的内容区间,遍历token过程中动态收集内容即可,完全不需要正则表达式:
- 用布尔变量
in_tag标记当前是否在标签内容区间内,初始为False - 用临时列表
current_content存储当前正在收集的标签内容,初始为空 - 遍历每个token时,遇到
#就切换状态,处于内容区间内的非#字符直接存入临时列表,遇到闭合#就把临时列表存入结果并清空
实现代码
tokens = ['0', '#', 'a', 'b', '#', '#', 'c', '#', '#', 'g', 'h', 'g', '#'] in_tag = False current_content = [] pair_list = [] for token in tokens: if token == '#': if in_tag: # 遇到闭合#,保存当前收集的内容 pair_list.append(current_content) current_content = [] in_tag = False else: # 遇到开启#,进入内容收集状态 in_tag = True else: if in_tag: current_content.append(token) # 输出验证 print(pair_list)
运行后输出结果和要求完全一致:[['a', 'b'], ['c'], ['g', 'h', 'g']]
业务场景适配验证
针对社交媒体哈希标签场景,比如句子I live in #United States# and #New York#!拆分后的token列表:
tokens = ['I', 'live', 'in', '#', 'United', 'States', '#', 'and', '#', 'New', 'York', '#', '!']
运行上述代码得到的结果为:[['United', 'States'], ['New', 'York']],符合业务提取需求。
内容的提问来源于stack exchange,提问作者marlon
相关产品推荐
相关产品推荐

