You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3合并多组残基计数列表并生成表格型数据的方法

Hey there! Let's break down your problem and work through a solid solution for processing those protein sequences, aggregating the data, and exporting it properly.

1. Current Input/Output Context

Input Sequence Format

Each input sequence follows this structure:

1A5ZA:A|PDBID|CHAIN|SQUENCEMKIGIVGLGRVGSSTAFAL

Generated Single-Sequence Data

Your GetResCount function produces lists like these for each individual sequence:

list1=[['name', '1A5ZA'], ['length', 83], ['A', 28], ['V', 31], ['I', 24]]
list2=[['name', '1AJ8A'], ['length', 49], ['A', 18], ['V', 11], ['I', 20]]
list3=[['name', '1AORA'], ['length', 96], ['A', 32], ['V', 49], ['I', 15]]

Desired Aggregated Format

You want to combine all these individual datasets into a unified structure like:

final=[['name', '1A5ZA', '1AJ8A', '1AORA'], ['length', 83, 49, 96], ['A', 28, 18, 32], ['V', 31, 11, 49], ['I', 24, 20, 15]]
2. Existing Function Notes

First, let's clean up the formatting of your current GetResCount function for readability:

def GetResCount(sequence):
    residues=[['A',0],['V',0],['I',0],['L',0],['M',0],['F',0],['Y',0],['W',0], 
              ['S',0],['T',0],['N',0],['Q',0],['C',0],['U',0],['G',0],['P',0],['R',0], 
              ['H',0],['K',0],['D',0],['E',0]]
    name=sequence[0:5]
    AAseq=sequence[27:]
    for AA in AAseq:
        for n in range(len(residues)):
            if residues[n][0] == AA:
                residues[n][1] = residues[n][1]+1
    length=len(AAseq)
    nameList=(['name', name])
    lengthList=(['length', length])
    residues.insert(0, lengthList)
    residues.insert(0, nameList)
    return residues

A quick optimization note: those nested loops for counting residues are a bit inefficient. Using a dictionary for counting will speed things up, especially with large datasets.

3. Core Requirements

To recap your needs clearly:

  • Extract the sequence name (first 5 characters of the input string)
  • Calculate the length of the amino acid (AA) sequence
  • Count occurrences of each predefined AA residue
  • Aggregate data from multiple sequences into a single, easy-to-use structure
  • Export the aggregated data to a table format for further use
4. Optimized Solution

Your desired list-of-lists format works, but a dictionary of lists is more flexible and maintainable. For example:

aggregated_data = {
    'name': ['1A5ZA', '1AJ8A', '1AORA'],
    'length': [83, 49, 96],
    'A': [28, 18, 32],
    'V': [31, 11, 49],
    'I': [24, 20, 15]
}

This structure lets you quickly access all values for a specific key (e.g., aggregated_data['A'] gives all alanine counts) and is easy to convert to your desired list-of-lists format or a table.

Updated Implementation Code

Let's rewrite the processing and aggregation steps for better efficiency and readability:

Step 1: Optimized Sequence Processing Function

def process_sequence(sequence):
    # Define all residues we need to count
    residue_keys = ['A', 'V', 'I', 'L', 'M', 'F', 'Y', 'W', 
                    'S', 'T', 'N', 'Q', 'C', 'U', 'G', 'P', 'R', 
                    'H', 'K', 'D', 'E']
    # Initialize counts to 0 using a dictionary
    residue_counts = {key: 0 for key in residue_keys}
    
    # Extract name and AA sequence from input
    name = sequence[:5]
    aa_seq = sequence[27:]
    length = len(aa_seq)
    
    # Count residues (faster than nested loops!)
    for aa in aa_seq:
        if aa in residue_counts:
            residue_counts[aa] += 1
    
    # Return all data for this sequence as a dictionary
    return {
        'name': name,
        'length': length,
        **residue_counts  # Unpack residue counts into the main dict
    }

Step 2: Aggregate Multiple Sequences

def aggregate_sequences(sequence_list):
    # Process all input sequences first
    processed_sequences = [process_sequence(seq) for seq in sequence_list]
    
    # Initialize aggregated data with empty lists for each key
    aggregated = {key: [] for key in processed_sequences[0].keys()}
    
    # Populate the aggregated lists with data from each sequence
    for seq_data in processed_sequences:
        for key, value in seq_data.items():
            aggregated[key].append(value)
    
    # Optional: Convert to your desired list-of-lists format
    list_of_lists = [[key] + values for key, values in aggregated.items()]
    
    return aggregated, list_of_lists

Step 3: Example Usage

# Sample input sequences
sample_sequences = [
    "1A5ZA:A|PDBID|CHAIN|SQUENCEMKIGIVGLGRVGSSTAFAL",
    "1AJ8A:A|PDBID|CHAIN|SQUENCEMKIGIVGLGRVGSSTA",
    "1AORA:A|PDBID|CHAIN|SQUENCEMKIGIVGLGRVGSSTAFALXYZ"
]

# Aggregate the data
aggregated_dict, aggregated_list = aggregate_sequences(sample_sequences)

# Print the list-of-lists format you requested
print(aggregated_list)
5. Export to Table Format

To export the aggregated data to a widely supported table format (like CSV), use Python's built-in csv module:

import csv

def export_to_csv(aggregated_dict, filename='residue_counts.csv'):
    # Define column order (name, length, then residues)
    fieldnames = ['name', 'length'] + [k for k in aggregated_dict.keys() if k not in ['name', 'length']]
    
    with open(filename, 'w', newline='') as f:
        writer = csv.DictWriter(f, fieldnames=fieldnames)
        writer.writeheader()
        
        # Write one row per sequence
        for i in range(len(aggregated_dict['name'])):
            row = {key: aggregated_dict[key][i] for key in fieldnames}
            writer.writerow(row)

# Usage
export_to_csv(aggregated_dict)

This creates a CSV file where each row represents a sequence, with columns for name, length, and each residue count—perfect for spreadsheets or further analysis.

Quick Notes on Your Original Approach

Your initial list-of-lists format is totally valid for simple use cases, but the dictionary approach is more maintainable (e.g., easy to add/remove residues, access specific data quickly). The optimized counting using a dictionary also speeds up processing, especially with large numbers of sequences.


内容的提问来源于stack exchange,提问作者Sam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:59:00