You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将指定格式数据集转为Python字典(去除指定关键词)

问题解决:将bag数据集转换为Python字典

问题描述

用户拥有如下格式的数据集:

bag 1:
weight: 9.4
value: 57
bag 2:
weight: 7.4
value: 94
...(其余bag条目略)

需要将其转换为类似{'1': [9.4, 57], '2': [7.4, 94]}的Python字典,同时去除“bag”“weight”“value”字样。但用户编写的代码无输出,代码如下:

capacity = 285
val = 0
weight = 0
b ='bag '
w = 'weight'
v = 'value'
Bag_dict = {}

with open('BankProblem.txt') as f:
    current_id = ''
    for line in f:
        if line.startswith('bag '):
            current_id = line.replace(':', '')
            if line.startswith('weight:'):
                Bag_dict.append((current_id, line))
                if line.startswith('value:'):
                    Bag_dict.append((current_id, weight, line))
print(Bag_dict)

原代码问题分析

  1. 逻辑嵌套错误:在判断行以bag 开头的分支里,又判断该行是否以weight:开头,这两个条件不可能同时成立,导致后续代码永远不会执行。
  2. 字典操作错误:Bag_dict是字典类型,没有append方法,字典需要通过键值对的方式添加数据。
  3. 数据提取不完整:没有从weight:和value:行中提取具体数值,只是保留了原始行文本,也没有转换为正确的数据类型。
  4. ID提取错误:line.replace(':', '')得到的是'bag 1',没有提取出纯数字作为字典的键。

修正后的代码

Bag_dict = {}

with open('BankProblem.txt') as f:
    current_id = None
    current_weight = None
    current_value = None
    for line in f:
        line = line.strip()  # 去除首尾空白(缩进、换行符等)
        if not line:
            continue  # 跳过空行
        
        if line.startswith('bag '):
            # 先把上一个bag的数据存入字典(如果存在)
            if current_id and current_weight is not None and current_value is not None:
                Bag_dict[current_id] = [current_weight, current_value]
            # 提取bag的数字ID
            current_id = line.replace('bag ', '').rstrip(':')
            # 重置当前bag的权重和价值
            current_weight = None
            current_value = None
        elif line.startswith('weight:'):
            # 提取权重数值并转为float类型
            current_weight = float(line.split(':')[1].strip())
        elif line.startswith('value:'):
            # 提取价值数值并转为int类型
            current_value = int(line.split(':')[1].strip())
    
    # 处理最后一个bag的数据
    if current_id and current_weight is not None and current_value is not None:
        Bag_dict[current_id] = [current_weight, current_value]

print(Bag_dict)

代码说明

  • 使用strip()处理每行的空白字符,避免缩进和换行影响判断逻辑。
  • 分阶段收集每个bag的ID、权重和价值,当一个bag的所有数据收集完成后,再存入字典。
  • 通过split(':')提取冒号后的数值部分,并转换为对应的数值类型(权重为float,价值为int)。
  • 循环结束后处理最后一个bag的数据,避免遗漏。
  • 字典通过键: 值的方式添加条目,符合字典的操作规范。

内容的提问来源于stack exchange,提问作者DasIstDread

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 19:30:54