You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

加载Feature-Barcode矩阵至Python时触发IndexError:列表索引越界

问题解决:读取10X Genomics Barcodes文件时的索引越界错误

错误原因

10X Genomics的barcodes.tsv.gz文件默认仅包含一列数据(细胞条码字符串),你的代码尝试访问row[2],但列表仅存在row[0]这一有效索引,因此触发IndexError。

另外代码存在低效问题:读取features文件时重复打开三次,可一次性读取所有数据再拆分,减少IO开销。

修正后的代码

import csv
import gzip
import os
import scipy.io
import pandas as pd
import json

matrix_dir = "filtered_feature_bc_matrix"
# 读取稀疏矩阵
mat = scipy.io.mmread(os.path.join(matrix_dir, "GSM4775593_Q5matrix.mtx.gz"))

# 一次性读取features文件,避免重复IO操作
features_path = os.path.join(matrix_dir, "GSM4775593_Q5genes.tsv.gz")
feature_data = []
with gzip.open(features_path, mode="rt") as f:
    reader = csv.reader(f, delimiter="\t")
    feature_data = list(reader)

feature_ids = [row[0] for row in feature_data]
gene_names = [row[1] for row in feature_data]
feature_types = [row[2] for row in feature_data]

# 读取barcodes文件,使用正确的索引row[0]
barcodes_path = os.path.join(matrix_dir, "GSM4775593_Q5barcodes.tsv.gz")
barcodes = []
with gzip.open(barcodes_path, mode="rt") as f:
    reader = csv.reader(f, delimiter="\t")
    barcodes = [row[0] for row in reader]

额外优化:直接转为DataFrame并保存为CSV(可选)

如果最终目标是生成CSV文件,可直接将稀疏矩阵转为DataFrame,简化后续处理:

# 将稀疏矩阵转为DataFrame
df = pd.DataFrame.sparse.from_spmatrix(mat, index=gene_names, columns=barcodes)
# 保存为CSV文件
df.to_csv("expression_matrix.csv")

内容的提问来源于stack exchange,提问作者fdzsfhaS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 08:25:02