基于条件从Python字典中提取特定范围元素的实现需求
Python嵌套字典页面提取:规则实现与特殊场景处理
需求明确
从嵌套字典ip_dict中提取页面,需严格遵循以下规则:
- 提取所有标记为
FP(首页)和LP(末页)的页面 - 仅保留位于某一个FP和其后续第一个LP之间的
Others页面,其余Others直接忽略 - 若某个FP之后没有对应的LP,则单独保留该FP
特殊场景:LP前置的处理逻辑
当LP出现在FP之前时,该LP没有对应的前置FP作为起始节点,因此不提取该LP,其前后的Others也一并忽略。后续出现的FP仍按正常规则处理:
- 示例输入片段:
[{'type': 'LP', 'content': '页1'}, {'type': 'Others', 'content': '页2'}, {'type': 'FP', 'content': '页3'}, {'type': 'Others', 'content': '页4'}] - 处理结果:仅保留FP(页3),因为该FP后无LP,符合规则3;前置的LP和对应Others全部被过滤
实现方案
代码思路
- 先从嵌套字典中提取页面序列(假设页面列表存储在
pages键下,可根据实际结构调整) - 遍历页面序列时,维护当前跟踪的FP节点:
- 遇到FP时,若之前存在未匹配LP的FP,先将其加入结果;再启动新的FP跟踪,开始收集后续内容
- 遇到LP时,若当前有正在跟踪的FP,则将LP加入当前组,完成该FP-LP序列的提取,结束跟踪
- 遇到Others时,仅在有正在跟踪的FP且未遇到LP的情况下,才将其加入当前组
- 遍历结束后,若仍有未匹配LP的FP,单独加入结果
代码实现
def extract_pages(ip_dict): # 从嵌套字典中获取页面列表,可根据实际结构修改键名 pages = ip_dict.get('pages', []) op_dict = {'extracted_pages': []} current_fp = None current_group = [] for page in pages: page_type = page.get('type') if page_type == 'FP': # 先处理上一个未匹配LP的FP if current_fp is not None: op_dict['extracted_pages'].append(current_fp) # 启动新的FP跟踪 current_fp = page current_group = [current_fp] elif page_type == 'LP': if current_fp is not None: # 完成当前FP-LP组的收集 current_group.append(page) op_dict['extracted_pages'].extend(current_group) current_fp = None current_group = [] # 无对应FP的LP直接忽略 elif page_type == 'Others': # 仅在跟踪FP且未遇到LP时收集Others if current_fp is not None: current_group.append(page) # 处理遍历结束后剩余的未匹配LP的FP if current_fp is not None: op_dict['extracted_pages'].append(current_fp) return op_dict
测试验证
案例1:正常FP-Others-LP序列
输入:
ip_dict = { 'pages': [ {'type': 'FP', 'content': '首页1'}, {'type': 'Others', 'content': '中间页1'}, {'type': 'Others', 'content': '中间页2'}, {'type': 'LP', 'content': '末页1'}, {'type': 'Others', 'content': '无关页'}, {'type': 'FP', 'content': '首页2'}, {'type': 'LP', 'content': '末页2'} ] }
输出:
{ 'extracted_pages': [ {'type': 'FP', 'content': '首页1'}, {'type': 'Others', 'content': '中间页1'}, {'type': 'Others', 'content': '中间页2'}, {'type': 'LP', 'content': '末页1'}, {'type': 'FP', 'content': '首页2'}, {'type': 'LP', 'content': '末页2'} ] }
案例2:FP后无LP
输入:
ip_dict = { 'pages': [ {'type': 'Others', 'content': '无关页'}, {'type': 'FP', 'content': '孤立首页'}, {'type': 'Others', 'content': '后续页'} ] }
输出:
{ 'extracted_pages': [{'type': 'FP', 'content': '孤立首页'}] }
案例3:LP前置场景
输入:
ip_dict = { 'pages': [ {'type': 'LP', 'content': '前置末页'}, {'type': 'Others', 'content': '前置中间页'}, {'type': 'FP', 'content': '有效首页'}, {'type': 'Others', 'content': '有效中间页'}, {'type': 'LP', 'content': '有效末页'} ] }
输出:
{ 'extracted_pages': [ {'type': 'FP', 'content': '有效首页'}, {'type': 'Others', 'content': '有效中间页'}, {'type': 'LP', 'content': '有效末页'} ] }
内容的提问来源于stack exchange,提问作者spectre
相关产品推荐
相关产品推荐

