You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

批量处理Ensembl基因表:识别特定字符组合并移除固定长度后缀

How to Batch Remove Fixed-Length Suffix from Gene Names

Since your gene names follow a consistent pattern where the unwanted suffix is exactly 16 characters long (e.g., _ENST00000WWWWWW), we can leverage this fixed length to trim the excess in bulk—no manual editing needed. Below are practical solutions using tools common in bioinformatics and data processing:

Python (Ideal for Large Datasets)

If you’re working with a CSV/Excel file, pandas makes this quick and scalable. Let’s assume your gene names are in a column labeled Gene:

import pandas as pd

# Load your gene table
df = pd.read_csv("your_gene_table.csv")

# Trim the last 16 characters from each gene name
df["Gene"] = df["Gene"].str[:-16]

# Save the cleaned table
df.to_csv("cleaned_gene_table.csv", index=False)

If you’re dealing with a list of gene names directly, a simple list comprehension works:

gene_names = ["GeneX_ENST00000123456", "GeneY_ENST00000789012"]
cleaned_names = [name[:-16] for name in gene_names]

R (Bioinformatics Standard)

Use base R or the stringr package for seamless processing:

Base R:

# Read your data
df <- read.csv("your_gene_table.csv")

# Remove the last 16 characters from the Gene column
df$Gene <- substr(df$Gene, 1, nchar(df$Gene) - 16)

# Export the cleaned table
write.csv(df, "cleaned_gene_table.csv", row.names = FALSE)

With stringr (more readable):

library(stringr)

df$Gene <- str_remove(df$Gene, ".{16}$")

Excel/Google Sheets (No Coding Required)

For a GUI-based approach, use this formula in the cell adjacent to your first gene name (e.g., if the gene name is in cell A1):

=LEFT(A1, LEN(A1)-16)

Drag the fill handle down to apply this to all rows, then copy-paste the values over the original column if you want to replace the raw data.

Command Line (Fast for Text/TSV Files)

If you’re working with plain text or TSV files, awk or sed handle large datasets efficiently:

Awk (for CSV, assuming gene name is the first column):

awk -F, '{sub(/.{16}$/,"",$1); print}' input.csv > output.csv

Sed (for any text file):

sed 's/.\{16\}$//' input.txt > output.txt

All these methods work perfectly even with duplicate gene names—you don’t need to adjust anything to account for repeats.

内容的提问来源于stack exchange,提问作者Jordan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:14:37