You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于公共列合并行数不等的多文件生成计数矩阵的Python代码调试

问题:基于公共列合并多文件生成计数矩阵(Python调试)

需求说明

需基于公共列(id列)合并行数不等的多文件,生成计数矩阵,缺失值用0填充。

输入文件示例

File1.txt

id  SRR1071717
chr1:15039:-::chr1:15795:-  2
chr1:15948:-::chr1:16606:-  6

File2.txt

id  SRR1079830
chr1:11672:+::chr1:12009:+  10
chr1:11845:+::chr1:12009:+  7
chrY:9756574:+::chrY:9757796:+  0

期望输出

id  SRR1071717 SRR1079830
chr1:15039:-::chr1:15795:- 2 0
chr1:15948:-::chr1:16606:- 6 0
chr1:11672:+::chr1:12009:+ 0 10
chr1:11845:+::chr1:12009:+ 0 7
chrY:9756574:+::chrY:9757796:+ 0 0

现有代码及问题

编写的Matrix.py代码

import sys

columns = []
data = {}
ids = set()
for filename in sys.argv[1:]:
    with open(filename, 'rU') as f:
        key = next(f).strip().split()[1]
        columns.append(key)
        data[key] = {}
        for line in f:
            if line.strip():
                id, value = line.strip().split()
                try:
                    data[key][int(id)] = value
                except ValueError as exc:
                    raise ValueError(
                        "Problem in line: '{}' '{}' '{}'".format(
                            id, value, line.rstrip()))

                ids.add(int(id))

print('\t'.join(['ID'] + columns))

for id in sorted(ids):
    line = []
    for column in columns:
        line.append(data[column].get(id, '0'))
    print('\t'.join([str(id)] + line))

运行结果异常

执行命令 python3 matrix.py File*.txt 后,仅输出两行,不符合预期:

id SRR1071717 SRR1079830
chrY:9756574:+::chrY:9757796:+ 0 0

问题分析与代码修正

核心错误

代码错误地将字符串类型的id强制转换为整数,而输入文件中的id是类似chr1:15039:-::chr1:15795:-的字符串标识符,无法转换为整数。这导致大部分id处理时抛出异常,最终仅残留少量未触发异常的记录。

修正后的代码

import sys

columns = []
data = {}
ids = set()
for filename in sys.argv[1:]:
    with open(filename, 'r') as f:
        # 读取表头,获取样本名
        header_line = next(f).strip()
        key = header_line.split()[1]
        columns.append(key)
        data[key] = {}
        for line_num, line in enumerate(f, start=2):
            line = line.strip()
            if not line:
                continue
            try:
                # 分割行,确保数值含空格也能正确处理
                id_str, value = line.split(maxsplit=1)
                # 直接用字符串id作为键,无需转换为int
                data[key][id_str] = value.strip()
                ids.add(id_str)
            except ValueError as exc:
                raise ValueError(
                    f"文件 {filename} 第 {line_num} 行格式错误: '{line}'\n错误原因: {exc}"
                )

# 输出表头,保持与期望一致的'id'
print('\t'.join(['id'] + columns))

# 按字典序排序id并输出
for id_str in sorted(ids):
    row = [id_str]
    for col in columns:
        # 缺失值填充为'0'
        row.append(data[col].get(id_str, '0'))
    print('\t'.join(row))

关键修改点

  1. 移除不必要的int转换:所有涉及int(id)的代码改为直接使用原始字符串id_str,匹配输入文件中id的字符串类型。
  2. 优化异常处理:扩大try块范围覆盖行分割和数据存储,增加行号提示便于定位格式错误。
  3. 修正表头输出:将['ID']改为['id'],与输入输出格式保持统一。
  4. 改进行分割逻辑:使用split(maxsplit=1),确保数值部分含空格时也能正确分割。
  5. 替换过时模式:将废弃的rU模式改为标准的r模式,适配Python 3环境。

测试验证

执行修正后的代码,将得到与期望完全一致的输出结果。

内容的提问来源于stack exchange,提问作者Shafaque Zahra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 01:13:24