基于Phobius输出在Python中提取序列区域及结果验证的技术咨询
Hey there! I totally get how tricky it can be to wrangle oddly formatted tool outputs when you're still building up your Python skills. Let's walk through exactly how to parse that Phobius output and extract those sequence regions you need.
Step 1: Break Down the Phobius Output Format
First, let's clarify what each part of the output means. Here's your sample output for reference:
SEQENCE ID TM SP PREDICTION
YOL154W_Q12512_Saccharomyces_cerevisiae 0 Y n8-15c20/21o
YDR481C_P11491_Saccharomyces_cerevisiae 1 0 i34-53o
YAL007C_P39704_Saccharomyces_cerevisiae 1 Y n5-20c25/26o181-207i
YAR028W_P39548_Saccharomyces_cerevisiae 2 0 i51-69o75-97i
YBL040C_P18414_Saccharomyces_cerevisiae 7 0 o6-26i38-56o62-80i101-119o125-143i15...
Each line (after the header) has 4 core components:
- Sequence ID: Unique protein identifier (e.g.,
YOL154W_Q12512_Saccharomyces_cerevisiae) - TM: Number of transmembrane domains
- SP: Signal peptide flag (
Y= has signal peptide,0/n= no signal peptide) - PREDICTION: Compact topology string where each segment starts with a letter (
n=non-cytoplasmic,c=cytoplasmic,i=transmembrane interior,o=transmembrane exterior) followed by a position range (e.g.,n8-15= non-cytoplasmic region from 8 to 15)
Step 2: Python Code to Extract Sequence Regions
I've written a simple, well-commented script that handles this parsing. It's designed to be easy to follow even if you're new to Python:
import re def parse_phobius_output(output_text): # Store parsed results for each sequence in a list parsed_results = [] # Split input text into lines, skip the first header line lines = output_text.strip().split('\n')[1:] for line in lines: # Skip empty lines if any if not line.strip(): continue # Split the line into individual fields line_parts = line.split() seq_id = line_parts[0] tm_count = int(line_parts[1]) # Convert TM count to integer sp_flag = line_parts[2] prediction_string = line_parts[3] # Use regex to extract each region from the prediction string # Pattern matches a letter (n/c/i/o) followed by position range(s) region_pattern = re.compile(r'([ncio])(\d+(?:[-/]\d+)?)') raw_regions = region_pattern.findall(prediction_string) # Clean up regions into a readable, structured format formatted_regions = [] for region_type, position in raw_regions: # Handle different position formats if '-' in position: start, end = map(int, position.split('-')) elif '/' in position: # For cases like 25/26, treat as start and end positions start, end = map(int, position.split('/')) else: # Single number case (rare, but handle it gracefully) start = end = int(position) formatted_regions.append({ 'type': region_type, 'start': start, 'end': end }) # Add parsed data for this sequence to results parsed_results.append({ 'sequence_id': seq_id, 'tm_domains': tm_count, 'signal_peptide': sp_flag, 'topology_regions': formatted_regions }) return parsed_results # -------------------------- # Example Usage # -------------------------- # Replace this with your actual Phobius output text sample_phobius_output = """SEQENCE ID TM SP PREDICTION YOL154W_Q12512_Saccharomyces_cerevisiae 0 Y n8-15c20/21o YDR481C_P11491_Saccharomyces_cerevisiae 1 0 i34-53o YAL007C_P39704_Saccharomyces_cerevisiae 1 Y n5-20c25/26o181-207i YAR028W_P39548_Saccharomyces_cerevisiae 2 0 i51-69o75-97i YBL040C_P18414_Saccharomyces_cerevisiae 7 0 o6-26i38-56o62-80i101-119o125-143i15...""" # Parse the output results = parse_phobius_output(sample_phobius_output) # Print results to verify for entry in results: print(f"Protein ID: {entry['sequence_id']}") print(f"Transmembrane Domains: {entry['tm_domains']}") print(f"Signal Peptide: {entry['signal_peptide']}") print("Topology Regions:") for region in entry['topology_regions']: print(f" - Type: {region['type']}, Position: {region['start']}-{region['end']}") print("-" * 50)
Step 3: Use the Code with a File
If your Phobius output is saved in a text file instead of a string, replace the example usage section with this code to read directly from the file:
# Read Phobius output from a file with open('phobius_results.txt', 'r') as file: phobius_output = file.read() # Parse the file content results = parse_phobius_output(phobius_output) # Now you can work with the parsed data (e.g., save to CSV, filter specific regions)
Quick Tips for Beginners
- The
remodule handles pattern matching, which is perfect for pulling structured data out of messy text. The regex pattern I used covers all common Phobius position formats. - All variables have descriptive names, and comments explain each step—take your time to read through them to understand what's happening.
- If you run into edge cases (like unexpected formats in the prediction string), you can tweak the regex pattern or add extra handling in the region formatting loop.
内容的提问来源于stack exchange,提问作者Fini

