如何将指定格式数据集转为Python字典(去除指定关键词)
问题解决:将bag数据集转换为Python字典
问题描述
用户拥有如下格式的数据集:
bag 1:
weight: 9.4
value: 57
bag 2:
weight: 7.4
value: 94
...(其余bag条目略)
需要将其转换为类似{'1': [9.4, 57], '2': [7.4, 94]}的Python字典,同时去除“bag”“weight”“value”字样。但用户编写的代码无输出,代码如下:
capacity = 285 val = 0 weight = 0 b ='bag ' w = 'weight' v = 'value' Bag_dict = {} with open('BankProblem.txt') as f: current_id = '' for line in f: if line.startswith('bag '): current_id = line.replace(':', '') if line.startswith('weight:'): Bag_dict.append((current_id, line)) if line.startswith('value:'): Bag_dict.append((current_id, weight, line)) print(Bag_dict)
原代码问题分析
- 逻辑嵌套错误:在判断行以
bag开头的分支里,又判断该行是否以weight:开头,这两个条件不可能同时成立,导致后续代码永远不会执行。 - 字典操作错误:
Bag_dict是字典类型,没有append方法,字典需要通过键值对的方式添加数据。 - 数据提取不完整:没有从
weight:和value:行中提取具体数值,只是保留了原始行文本,也没有转换为正确的数据类型。 - ID提取错误:
line.replace(':', '')得到的是'bag 1',没有提取出纯数字作为字典的键。
修正后的代码
Bag_dict = {} with open('BankProblem.txt') as f: current_id = None current_weight = None current_value = None for line in f: line = line.strip() # 去除首尾空白(缩进、换行符等) if not line: continue # 跳过空行 if line.startswith('bag '): # 先把上一个bag的数据存入字典(如果存在) if current_id and current_weight is not None and current_value is not None: Bag_dict[current_id] = [current_weight, current_value] # 提取bag的数字ID current_id = line.replace('bag ', '').rstrip(':') # 重置当前bag的权重和价值 current_weight = None current_value = None elif line.startswith('weight:'): # 提取权重数值并转为float类型 current_weight = float(line.split(':')[1].strip()) elif line.startswith('value:'): # 提取价值数值并转为int类型 current_value = int(line.split(':')[1].strip()) # 处理最后一个bag的数据 if current_id and current_weight is not None and current_value is not None: Bag_dict[current_id] = [current_weight, current_value] print(Bag_dict)
代码说明
- 使用
strip()处理每行的空白字符,避免缩进和换行影响判断逻辑。 - 分阶段收集每个bag的ID、权重和价值,当一个bag的所有数据收集完成后,再存入字典。
- 通过
split(':')提取冒号后的数值部分,并转换为对应的数值类型(权重为float,价值为int)。 - 循环结束后处理最后一个bag的数据,避免遗漏。
- 字典通过
键: 值的方式添加条目,符合字典的操作规范。
内容的提问来源于stack exchange,提问作者DasIstDread
相关产品推荐
相关产品推荐

