You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python列表元素排序异常:组装报告信息提取位置指定问题求助

Fixing the Element Order in Your Assembly Report Extraction

The issue here is that you're appending elements to the description list in the order they appear in the file, but your report has the Assembly name line first. That's why your output list ends up with assembly at the top instead of the desired ["species", "strain", "assembly"] order.

To fix this, we can use a temporary dictionary to store each value by its key, then build the final list using your desired order—this way, the file's line order won't affect the result. Here's how to adjust your code:

import os

report_dict = {}
assembly_report_dir = "your_directory_path"  # Make sure this points to your actual directory

for root, dirs, reports in os.walk(assembly_report_dir):
    for report in reports:
        # Extract accession (kept your original logic here)
        accession = '_'.join(report.strip().split('/')[-1].replace('_assembly_report.txt', '').split('_')[0:2])
        path = os.path.join(root, report)  # Use root to handle subdirectories correctly
        
        # Initialize a dictionary to hold each element with clear keys
        desc_elements = {
            "species": None,
            "strain": None,
            "assembly": None
        }
        
        with open(path, 'r') as inputfile:
            for line in inputfile.readlines():
                line = line.strip()
                if line.startswith('Organism name: '):
                    organism = line.split(': ')[-1].split(' (', 1)[0]
                    desc_elements["species"] = ' '.join(organism.split(' ')[:2])
                elif line.startswith('Infraspecific name: strain='):
                    # Removed the unnecessary space-stripping to keep "BPIC 631" intact
                    desc_elements["strain"] = line.split('strain=')[-1].strip()
                elif line.startswith('Assembly name: '):
                    desc_elements["assembly"] = line.split(': ')[-1]
        
        # Build the final list explicitly in your preferred order
        description = [desc_elements["species"], desc_elements["strain"], desc_elements["assembly"]]
        report_dict[accession] = description

print(report_dict)

Key Changes Explained:

  1. Temporary Dictionary: Using desc_elements lets us store each value with a clear label, so we don't depend on the order of lines in the report file.
  2. Explicit List Order: After collecting all values, we construct the description list exactly in the sequence you want: species → strain → assembly.
  3. Minor Cleanup: Removed the .replace(' ','') from strain extraction—this was turning "BPIC 631" into "BPIC631", which likely wasn't intentional. If you do need to strip spaces, you can add that back.

Testing this with your sample report will produce the correct output list:
["Pseudomonas avellanae", "BPIC 631", "Pav631_1.0"]

内容的提问来源于stack exchange,提问作者F.Lira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:53:59