基于Python实现依据分类文件批量重命名生物信息学bin文件
批量重命名Bin文件为分类学属名的Python实现
我来帮你搞定这个生物信息学数据整理的需求,用Python实现完全没问题,下面是具体的思路和可运行的代码:
问题拆解
咱们的核心需求可以拆成这几步:
- 遍历
binned/下的所有日期命名子文件夹(比如90-20-09-2018、90-25-04-2018) - 每个子文件夹里,从
bins/quality/taxonomy.txt读取Bin文件和属名的对应关系 - 把
bins/下的原始Bin文件(比如90-20-09-2018.001)重命名为对应的属名(比如Lactobacillus) - 处理特殊情况:比如有的分类条目没有属级信息(只到科)、多个Bin对应同一个属名的冲突
实现思路
- 路径遍历:用Python的
pathlib模块(比传统os模块更直观)来遍历目标文件夹 - 解析分类文件:读取
taxonomy.txt,跳过表头行,拆分每行数据,提取Bin Id和g__后的属名部分 - 重命名逻辑:根据解析出的映射关系,逐个重命名Bin文件;遇到重名时自动添加序号后缀(比如
Lactobacillus_1) - 异常处理:处理文件不存在、权限不足等常见问题,避免程序崩溃
完整代码
from pathlib import Path from collections import defaultdict def rename_bins_to_genus(root_dir: str = "binned"): root_path = Path(root_dir) # 遍历每个日期子文件夹 for date_folder in root_path.iterdir(): if not date_folder.is_dir(): continue # 定义关键路径 bins_dir = date_folder / "bins" taxonomy_file = bins_dir / "quality" / "taxonomy.txt" # 检查分类文件是否存在 if not taxonomy_file.exists(): print(f"警告:未找到分类文件 {taxonomy_file},跳过该文件夹") continue # 解析taxonomy.txt,构建Bin Id到属名的映射 bin_to_genus = {} # 统计每个属名出现的次数,处理重名 genus_count = defaultdict(int) with open(taxonomy_file, "r") as f: # 跳过表头行(直到找到包含"Bin Id"的行) for line in f: line = line.strip() if not line or line.startswith("---") or "Bin Id" not in line: continue # 读取表头后的有效数据行 break # 处理数据行 for line in f: line = line.strip() if not line or line.startswith("---"): continue # 拆分每行数据(按空格拆分,Taxonomy字段用分号分隔无空格) parts = line.split(maxsplit=3) if len(parts) < 4: print(f"警告:无效行,跳过:{line}") continue bin_id, _, _, taxonomy = parts # 提取属名(g__开头的部分) genus = None for taxon in taxonomy.split(";"): if taxon.startswith("g__"): genus = taxon[3:] # 去掉"g__"前缀 break # 如果没有属名,保留原Bin Id if not genus: genus = bin_id print(f"注意:Bin {bin_id} 没有属级分类信息,保留原名称") bin_to_genus[bin_id] = genus # 遍历bins目录下的所有Bin文件,进行重命名 for bin_file in bins_dir.iterdir(): if bin_file.is_dir() or bin_file.name not in bin_to_genus: continue # 跳过文件夹和不在映射里的文件 target_genus = bin_to_genus[bin_file.name] # 处理重名:如果该属名已经用过,添加序号 genus_count[target_genus] += 1 if genus_count[target_genus] > 1: new_name = f"{target_genus}_{genus_count[target_genus]-1}" else: new_name = target_genus new_path = bins_dir / new_name # 避免覆盖已存在的文件(如果有同名文件,自动递增序号) counter = 1 while new_path.exists(): new_name = f"{target_genus}_{counter}" new_path = bins_dir / new_name counter += 1 # 执行重命名 try: bin_file.rename(new_path) print(f"已重命名:{bin_file.name} -> {new_name}") except Exception as e: print(f"重命名失败 {bin_file.name}: {str(e)}") if __name__ == "__main__": rename_bins_to_genus()
注意事项
- 备份原始数据:运行代码前建议先备份
binned文件夹,避免意外数据丢失 - 处理无属名的条目:代码中如果遇到没有
g__的分类条目,会保留原Bin文件名,你可以根据需求修改这部分逻辑(比如改成科名或者自定义标记) - 重名处理:如果多个Bin对应同一个属名,代码会自动添加序号后缀(比如
Lactobacillus、Lactobacillus_1),避免文件覆盖 - 路径适配:如果你的
binned文件夹不在当前工作目录,可以修改rename_bins_to_genus函数的root_dir参数,传入完整路径
内容的提问来源于stack exchange,提问作者mortalknight55hotmailcom
相关产品推荐
相关产品推荐

