Python列表元素排序异常:组装报告信息提取位置指定问题求助
Fixing the Element Order in Your Assembly Report Extraction
The issue here is that you're appending elements to the description list in the order they appear in the file, but your report has the Assembly name line first. That's why your output list ends up with assembly at the top instead of the desired ["species", "strain", "assembly"] order.
To fix this, we can use a temporary dictionary to store each value by its key, then build the final list using your desired order—this way, the file's line order won't affect the result. Here's how to adjust your code:
import os report_dict = {} assembly_report_dir = "your_directory_path" # Make sure this points to your actual directory for root, dirs, reports in os.walk(assembly_report_dir): for report in reports: # Extract accession (kept your original logic here) accession = '_'.join(report.strip().split('/')[-1].replace('_assembly_report.txt', '').split('_')[0:2]) path = os.path.join(root, report) # Use root to handle subdirectories correctly # Initialize a dictionary to hold each element with clear keys desc_elements = { "species": None, "strain": None, "assembly": None } with open(path, 'r') as inputfile: for line in inputfile.readlines(): line = line.strip() if line.startswith('Organism name: '): organism = line.split(': ')[-1].split(' (', 1)[0] desc_elements["species"] = ' '.join(organism.split(' ')[:2]) elif line.startswith('Infraspecific name: strain='): # Removed the unnecessary space-stripping to keep "BPIC 631" intact desc_elements["strain"] = line.split('strain=')[-1].strip() elif line.startswith('Assembly name: '): desc_elements["assembly"] = line.split(': ')[-1] # Build the final list explicitly in your preferred order description = [desc_elements["species"], desc_elements["strain"], desc_elements["assembly"]] report_dict[accession] = description print(report_dict)
Key Changes Explained:
- Temporary Dictionary: Using
desc_elementslets us store each value with a clear label, so we don't depend on the order of lines in the report file. - Explicit List Order: After collecting all values, we construct the
descriptionlist exactly in the sequence you want:species→strain→assembly. - Minor Cleanup: Removed the
.replace(' ','')from strain extraction—this was turning "BPIC 631" into "BPIC631", which likely wasn't intentional. If you do need to strip spaces, you can add that back.
Testing this with your sample report will produce the correct output list:["Pseudomonas avellanae", "BPIC 631", "Pav631_1.0"]
内容的提问来源于stack exchange,提问作者F.Lira
相关产品推荐
相关产品推荐

