You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用HGNC命名规范批量替换基因矩阵的Gene ID(大矩阵适配)

Got it, let's tackle this problem efficiently—especially since you're dealing with a massive expression matrix that needs parallel processing. Here are two robust approaches using Python (with Dask for out-of-core parallelism) and R (with data.table + future for parallel chunking) to map your gene IDs to HGNC symbols:

Python Approach with Dask (Out-of-Core Parallelism)

Dask is perfect for large datasets that don't fit into memory—it automatically splits data into chunks and processes them in parallel without you having to manage threads manually.

Step-by-Step Code:

import dask.dataframe as dd
from dask.diagnostics import ProgressBar

# 1. Load the annotation table and create a gene-to-HGNC mapping
# Adjust sep to match your actual file delimiter (space/tab/comma)
annotations = dd.read_csv(
    "annotations.txt",
    sep="\s+",
    usecols=["gene", "hgnc_symbol"]
).set_index("gene")  # Set gene ID as index for fast joins

# 2. Load the large expression matrix
# Set 'gene ID' as the index to align with the annotation table
expr_matrix = dd.read_csv(
    "expression_matrix.txt",
    sep="\s+",
    index_col="gene ID"
)

# 3. Left join to map gene IDs to HGNC symbols
# This preserves all rows from your expression matrix, even if no HGNC match exists
merged = expr_matrix.join(annotations, how="left")

# 4. Handle unannotated genes (fill with original gene ID if no match)
merged["hgnc_symbol"] = merged["hgnc_symbol"].fillna(merged.index)

# 5. Set HGNC symbol as the new index and drop the old one
merged = merged.set_index("hgnc_symbol")

# 6. Save the result (single file for convenience)
with ProgressBar():
    merged.to_csv("expression_matrix_hgnc.csv", sep="\t", single_file=True)

Key Notes:

  • Dask handles parallelism under the hood—no need to write explicit thread/process code.
  • Adjust sep to match your actual file delimiters (space, tab, etc.).
  • For unannotated genes, we keep the original ID; you can replace fillna(merged.index) with fillna("UNANNOTATED") if you prefer labeling them instead.

R Approach with data.table + future (Parallel Chunking)

data.table is blazingly fast for large tabular data, and future lets us parallelize chunked processing to handle matrices too big for memory.

Step-by-Step Code:

library(data.table)
library(future)
library(furrr)

# 1. Set up parallel processing (uses all available cores by default)
plan(multisession, workers = availableCores())

# 2. Load annotation table and create a gene-to-HGNC mapping dictionary
annotations <- fread("annotations.txt", select = c("gene", "hgnc_symbol"))
gene_map <- setNames(annotations$hgnc_symbol, annotations$gene)

# 3. Load the large expression matrix
expr_matrix <- fread("expression_matrix.txt")

# 4. Define a function to replace gene IDs in a chunk
replace_gene_ids <- function(chunk) {
  # Replace matched IDs, keep original for unannotated genes
  chunk[, `gene ID` := fifelse(`gene ID` %in% names(gene_map), 
                               gene_map[`gene ID`], 
                               as.character(`gene ID`))]
  return(chunk)
}

# 5. Split matrix into chunks (adjust chunk_size based on your memory)
chunk_size <- 10000  # Tweak this to fit your system's RAM
chunks <- split(expr_matrix, ceiling(seq_len(nrow(expr_matrix)) / chunk_size))

# 6. Process chunks in parallel
result_chunks <- future_map(chunks, replace_gene_ids)

# 7. Combine chunks and save the final matrix
result <- rbindlist(result_chunks)
fwrite(result, "expression_matrix_hgnc.txt", sep = "\t")

Key Notes:

  • fread and fwrite are data.table's ultra-fast alternatives to base R's read/write functions.
  • Adjust chunk_size based on how much memory you have—smaller chunks use less RAM but take slightly longer.
  • The fifelse function efficiently replaces values without looping, which is critical for speed with large data.

内容的提问来源于stack exchange,提问作者Fernando Delgado Chaves

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:34:55