Python 2.7正则表达式:匹配可变数量的染色体位置表达式
Solution for Parsing Chromosome Location Strings in Python 2.7
Got it, let's work through this for your bioinformatics project. The goal is to extract those position numbers from your location strings and format them exactly as you need. Here's a straightforward regex-based solution tailored for Python 2.7:
Step 1: Regex Pattern Breakdown
First, let's define a regex that matches your input format and captures the numbers we care about:
r'NC_\d+\.\d+:g\.(\d+)(?:_(\d+))?'
Let's break this down clearly:
NC_\d+\.\d+:g\.: Matches the fixed prefix (likeNC_000023.11:g.). We escape dots with\because they’re special regex characters.(\d+): The first capture group, which grabs the mandatory first position number.(?:_(\d+))?: An optional non-capturing group (the?:makes it non-capturing) that matches an underscore followed by a second number. The?at the end means this part can appear 0 or 1 times. Inside it,(\d+)is the second capture group for the optional number.
Step 2: Python 2.7 Code Implementation
Since Python 2.7 doesn’t support f-strings, we’ll use str.format() instead to build the output. Here’s a complete working example:
import re # Your example location strings location_list = [ "NC_000023.11:g.154532082", "NC_000023.11:g.154532058_154532060", "NC_000023.11:g.154532046" ] # Precompile the regex for better performance with multiple strings location_pattern = re.compile(r'NC_\d+\.\d+:g\.(\d+)(?:_(\d+))?') for loc_str in location_list: match_result = location_pattern.match(loc_str) if match_result: first_pos = match_result.group(1) second_pos = match_result.group(2) # Generate the desired output format if second_pos: print("([{0}], [{1}])".format(first_pos, second_pos)) else: print("([{0}])".format(first_pos))
Step 3: Expected Output
When you run this code, you’ll get exactly the output you requested:
([154532082])
([154532058], [154532060])
([154532046])
Quick Notes
- This regex assumes all your input strings follow the exact format you provided. If you have edge cases (like different prefixes or more than two numbers), we can tweak the pattern easily.
- Precompiling the regex with
re.compile()is a good habit if you’re parsing a large number of strings—it speeds up the matching process.
内容的提问来源于stack exchange,提问作者Nulnulzeven
相关产品推荐
相关产品推荐

