Python 3合并多组残基计数列表并生成表格型数据的方法
Hey there! Let's break down your problem and work through a solid solution for processing those protein sequences, aggregating the data, and exporting it properly.
Input Sequence Format
Each input sequence follows this structure:
1A5ZA:A|PDBID|CHAIN|SQUENCEMKIGIVGLGRVGSSTAFAL
Generated Single-Sequence Data
Your GetResCount function produces lists like these for each individual sequence:
list1=[['name', '1A5ZA'], ['length', 83], ['A', 28], ['V', 31], ['I', 24]] list2=[['name', '1AJ8A'], ['length', 49], ['A', 18], ['V', 11], ['I', 20]] list3=[['name', '1AORA'], ['length', 96], ['A', 32], ['V', 49], ['I', 15]]
Desired Aggregated Format
You want to combine all these individual datasets into a unified structure like:
final=[['name', '1A5ZA', '1AJ8A', '1AORA'], ['length', 83, 49, 96], ['A', 28, 18, 32], ['V', 31, 11, 49], ['I', 24, 20, 15]]
First, let's clean up the formatting of your current GetResCount function for readability:
def GetResCount(sequence): residues=[['A',0],['V',0],['I',0],['L',0],['M',0],['F',0],['Y',0],['W',0], ['S',0],['T',0],['N',0],['Q',0],['C',0],['U',0],['G',0],['P',0],['R',0], ['H',0],['K',0],['D',0],['E',0]] name=sequence[0:5] AAseq=sequence[27:] for AA in AAseq: for n in range(len(residues)): if residues[n][0] == AA: residues[n][1] = residues[n][1]+1 length=len(AAseq) nameList=(['name', name]) lengthList=(['length', length]) residues.insert(0, lengthList) residues.insert(0, nameList) return residues
A quick optimization note: those nested loops for counting residues are a bit inefficient. Using a dictionary for counting will speed things up, especially with large datasets.
To recap your needs clearly:
- Extract the sequence name (first 5 characters of the input string)
- Calculate the length of the amino acid (AA) sequence
- Count occurrences of each predefined AA residue
- Aggregate data from multiple sequences into a single, easy-to-use structure
- Export the aggregated data to a table format for further use
Recommended Data Structure
Your desired list-of-lists format works, but a dictionary of lists is more flexible and maintainable. For example:
aggregated_data = { 'name': ['1A5ZA', '1AJ8A', '1AORA'], 'length': [83, 49, 96], 'A': [28, 18, 32], 'V': [31, 11, 49], 'I': [24, 20, 15] }
This structure lets you quickly access all values for a specific key (e.g., aggregated_data['A'] gives all alanine counts) and is easy to convert to your desired list-of-lists format or a table.
Updated Implementation Code
Let's rewrite the processing and aggregation steps for better efficiency and readability:
Step 1: Optimized Sequence Processing Function
def process_sequence(sequence): # Define all residues we need to count residue_keys = ['A', 'V', 'I', 'L', 'M', 'F', 'Y', 'W', 'S', 'T', 'N', 'Q', 'C', 'U', 'G', 'P', 'R', 'H', 'K', 'D', 'E'] # Initialize counts to 0 using a dictionary residue_counts = {key: 0 for key in residue_keys} # Extract name and AA sequence from input name = sequence[:5] aa_seq = sequence[27:] length = len(aa_seq) # Count residues (faster than nested loops!) for aa in aa_seq: if aa in residue_counts: residue_counts[aa] += 1 # Return all data for this sequence as a dictionary return { 'name': name, 'length': length, **residue_counts # Unpack residue counts into the main dict }
Step 2: Aggregate Multiple Sequences
def aggregate_sequences(sequence_list): # Process all input sequences first processed_sequences = [process_sequence(seq) for seq in sequence_list] # Initialize aggregated data with empty lists for each key aggregated = {key: [] for key in processed_sequences[0].keys()} # Populate the aggregated lists with data from each sequence for seq_data in processed_sequences: for key, value in seq_data.items(): aggregated[key].append(value) # Optional: Convert to your desired list-of-lists format list_of_lists = [[key] + values for key, values in aggregated.items()] return aggregated, list_of_lists
Step 3: Example Usage
# Sample input sequences sample_sequences = [ "1A5ZA:A|PDBID|CHAIN|SQUENCEMKIGIVGLGRVGSSTAFAL", "1AJ8A:A|PDBID|CHAIN|SQUENCEMKIGIVGLGRVGSSTA", "1AORA:A|PDBID|CHAIN|SQUENCEMKIGIVGLGRVGSSTAFALXYZ" ] # Aggregate the data aggregated_dict, aggregated_list = aggregate_sequences(sample_sequences) # Print the list-of-lists format you requested print(aggregated_list)
To export the aggregated data to a widely supported table format (like CSV), use Python's built-in csv module:
import csv def export_to_csv(aggregated_dict, filename='residue_counts.csv'): # Define column order (name, length, then residues) fieldnames = ['name', 'length'] + [k for k in aggregated_dict.keys() if k not in ['name', 'length']] with open(filename, 'w', newline='') as f: writer = csv.DictWriter(f, fieldnames=fieldnames) writer.writeheader() # Write one row per sequence for i in range(len(aggregated_dict['name'])): row = {key: aggregated_dict[key][i] for key in fieldnames} writer.writerow(row) # Usage export_to_csv(aggregated_dict)
This creates a CSV file where each row represents a sequence, with columns for name, length, and each residue count—perfect for spreadsheets or further analysis.
Your initial list-of-lists format is totally valid for simple use cases, but the dictionary approach is more maintainable (e.g., easy to add/remove residues, access specific data quickly). The optimized counting using a dictionary also speeds up processing, especially with large numbers of sequences.
内容的提问来源于stack exchange,提问作者Sam

