You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CSV数据解析需求:将TABLE格式数据转换为Python字典

解析CSV格式的数据馈送为Python字典

需求说明

我会定期下载一份数据馈送并保存为CSV文件,内容格式如下:

TABLE # 196712 / 9000_
>= 10   : 0.002
>= 5    : 0.001
>= 2    : 0.0005
>= 1    : 0.0002
>= 0.5  : 0.0001
>= 0.2  : 0.0001
>= 0.1  : 0.0001
>= 0.0001   : 0.0001
TABLE # 196714 / Dark
>= 0.0001   : 5e-05
TABLE # 196715 / GBD
>= 25   : 0.01
>= 10   : 0.005
>= 5    : 0.0025
>= 0.1  : 0.001
>= 0.0005   : 0.005

需要解析该文件,将数据整理为Python字典:以TABLE #后的数字作为唯一键(dict key),将后续以>=开头的行转换为(容量,惩罚值)的元组列表。预期输出格式如下:

{196712: [(10,0.002),(5,0.001),(2,0.0005),(1,0.0002),(0.5,0.0001),(0.2,0.0001),(0.1,0.0001),(0.0001, 0.0001)], 
 196714: [(0.0001,5e-05)], 
 196715: [(25,0.01),(10,0.005),(5,0.0025),(0.1,0.001),(0.0005,0.005)]}

如果不用Python,原本想用grep过滤行,但ID间行数不固定增加了复杂度,也希望能推荐更合适的数据结构。

Python实现方案

代码实现

def parse_data_feed(file_path):
    result = {}
    current_key = None
    current_list = []
    
    with open(file_path, 'r') as f:
        for line in f:
            line = line.strip()
            if not line:
                continue  # 跳过空行
            
            if line.startswith('TABLE #'):
                # 提取TABLE后的数字作为key
                parts = line.split()
                current_key = int(parts[2])
                current_list = []
                result[current_key] = current_list
            elif line.startswith('>='):
                # 解析容量和惩罚值
                parts = line.split(':')
                capacity_part = parts[0].replace('>=', '').strip()
                penalty_part = parts[1].strip()
                
                # 转换为数值类型(整数或浮点数)
                try:
                    capacity = int(capacity_part)
                except ValueError:
                    capacity = float(capacity_part)
                
                penalty = float(penalty_part)
                current_list.append((capacity, penalty))
    
    return result

# 使用示例
parsed_dict = parse_data_feed('your_data_file.csv')
print(parsed_dict)

代码说明

  • 逐行读取文件,跳过空行
  • 遇到TABLE #开头的行,提取数字作为字典的键,并初始化对应的空列表
  • 遇到>=开头的行,拆分出容量和惩罚值,转换为合适的数值类型(整数或浮点数),添加到当前列表中
  • 最终返回整理好的字典

非Python工具方案

如果不想用Python,可以用awk处理,它能轻松处理行与行之间的关联逻辑,无需担心ID间行数不固定的问题:

BEGIN { current_key = "" }
/^TABLE #/ {
    current_key = $3
    next
}
/^>=/ {
    split($0, parts, ":")
    gsub(">=", "", parts[1])
    capacity = parts[1] + 0  # 自动转换为数值
    penalty = parts[2] + 0
    print current_key, capacity, penalty
}

运行命令:awk -f parse.awk your_data_file.csv,输出会按每个TABLE ID对应的数据行打印,之后可以根据需要进一步整理成JSON等格式(比如结合jq)。

数据结构推荐

  • 如果后续需要频繁按TABLE ID查询,或者要对每个ID下的元组进行排序、查找操作,当前的字典+列表结构已经很合适:字典保证ID的唯一性和O(1)查询效率,列表保留了数据的原始顺序。
  • 如果需要更复杂的操作(比如按容量范围快速查找),可以考虑将每个ID下的列表转换为有序字典或者直接构建一个嵌套的字典(以容量为键,惩罚值为值),但要注意容量可能存在重复的情况(如果数据里有重复容量的话)。

内容的提问来源于stack exchange,提问作者ThatQuantDude

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 05:40:21