You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Phobius输出在Python中提取序列区域及结果验证的技术咨询

Hey there! I totally get how tricky it can be to wrangle oddly formatted tool outputs when you're still building up your Python skills. Let's walk through exactly how to parse that Phobius output and extract those sequence regions you need.

Step 1: Break Down the Phobius Output Format

First, let's clarify what each part of the output means. Here's your sample output for reference:

SEQENCE ID TM SP PREDICTION
YOL154W_Q12512_Saccharomyces_cerevisiae 0 Y n8-15c20/21o
YDR481C_P11491_Saccharomyces_cerevisiae 1 0 i34-53o
YAL007C_P39704_Saccharomyces_cerevisiae 1 Y n5-20c25/26o181-207i
YAR028W_P39548_Saccharomyces_cerevisiae 2 0 i51-69o75-97i
YBL040C_P18414_Saccharomyces_cerevisiae 7 0 o6-26i38-56o62-80i101-119o125-143i15...

Each line (after the header) has 4 core components:

  • Sequence ID: Unique protein identifier (e.g., YOL154W_Q12512_Saccharomyces_cerevisiae)
  • TM: Number of transmembrane domains
  • SP: Signal peptide flag (Y = has signal peptide, 0/n = no signal peptide)
  • PREDICTION: Compact topology string where each segment starts with a letter (n=non-cytoplasmic, c=cytoplasmic, i=transmembrane interior, o=transmembrane exterior) followed by a position range (e.g., n8-15 = non-cytoplasmic region from 8 to 15)

Step 2: Python Code to Extract Sequence Regions

I've written a simple, well-commented script that handles this parsing. It's designed to be easy to follow even if you're new to Python:

import re

def parse_phobius_output(output_text):
    # Store parsed results for each sequence in a list
    parsed_results = []
    
    # Split input text into lines, skip the first header line
    lines = output_text.strip().split('\n')[1:]
    
    for line in lines:
        # Skip empty lines if any
        if not line.strip():
            continue
        
        # Split the line into individual fields
        line_parts = line.split()
        seq_id = line_parts[0]
        tm_count = int(line_parts[1])  # Convert TM count to integer
        sp_flag = line_parts[2]
        prediction_string = line_parts[3]
        
        # Use regex to extract each region from the prediction string
        # Pattern matches a letter (n/c/i/o) followed by position range(s)
        region_pattern = re.compile(r'([ncio])(\d+(?:[-/]\d+)?)')
        raw_regions = region_pattern.findall(prediction_string)
        
        # Clean up regions into a readable, structured format
        formatted_regions = []
        for region_type, position in raw_regions:
            # Handle different position formats
            if '-' in position:
                start, end = map(int, position.split('-'))
            elif '/' in position:
                # For cases like 25/26, treat as start and end positions
                start, end = map(int, position.split('/'))
            else:
                # Single number case (rare, but handle it gracefully)
                start = end = int(position)
            
            formatted_regions.append({
                'type': region_type,
                'start': start,
                'end': end
            })
        
        # Add parsed data for this sequence to results
        parsed_results.append({
            'sequence_id': seq_id,
            'tm_domains': tm_count,
            'signal_peptide': sp_flag,
            'topology_regions': formatted_regions
        })
    
    return parsed_results

# --------------------------
# Example Usage
# --------------------------
# Replace this with your actual Phobius output text
sample_phobius_output = """SEQENCE ID TM SP PREDICTION
YOL154W_Q12512_Saccharomyces_cerevisiae 0 Y n8-15c20/21o
YDR481C_P11491_Saccharomyces_cerevisiae 1 0 i34-53o
YAL007C_P39704_Saccharomyces_cerevisiae 1 Y n5-20c25/26o181-207i
YAR028W_P39548_Saccharomyces_cerevisiae 2 0 i51-69o75-97i
YBL040C_P18414_Saccharomyces_cerevisiae 7 0 o6-26i38-56o62-80i101-119o125-143i15..."""

# Parse the output
results = parse_phobius_output(sample_phobius_output)

# Print results to verify
for entry in results:
    print(f"Protein ID: {entry['sequence_id']}")
    print(f"Transmembrane Domains: {entry['tm_domains']}")
    print(f"Signal Peptide: {entry['signal_peptide']}")
    print("Topology Regions:")
    for region in entry['topology_regions']:
        print(f"  - Type: {region['type']}, Position: {region['start']}-{region['end']}")
    print("-" * 50)

Step 3: Use the Code with a File

If your Phobius output is saved in a text file instead of a string, replace the example usage section with this code to read directly from the file:

# Read Phobius output from a file
with open('phobius_results.txt', 'r') as file:
    phobius_output = file.read()

# Parse the file content
results = parse_phobius_output(phobius_output)

# Now you can work with the parsed data (e.g., save to CSV, filter specific regions)

Quick Tips for Beginners

  • The re module handles pattern matching, which is perfect for pulling structured data out of messy text. The regex pattern I used covers all common Phobius position formats.
  • All variables have descriptive names, and comments explain each step—take your time to read through them to understand what's happening.
  • If you run into edge cases (like unexpected formats in the prediction string), you can tweak the regex pattern or add extra handling in the region formatting loop.

内容的提问来源于stack exchange,提问作者Fini

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:13:10