You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 2.7正则表达式:匹配可变数量的染色体位置表达式

Solution for Parsing Chromosome Location Strings in Python 2.7

Got it, let's work through this for your bioinformatics project. The goal is to extract those position numbers from your location strings and format them exactly as you need. Here's a straightforward regex-based solution tailored for Python 2.7:

Step 1: Regex Pattern Breakdown

First, let's define a regex that matches your input format and captures the numbers we care about:

r'NC_\d+\.\d+:g\.(\d+)(?:_(\d+))?'

Let's break this down clearly:

  • NC_\d+\.\d+:g\.: Matches the fixed prefix (like NC_000023.11:g.). We escape dots with \ because they’re special regex characters.
  • (\d+): The first capture group, which grabs the mandatory first position number.
  • (?:_(\d+))?: An optional non-capturing group (the ?: makes it non-capturing) that matches an underscore followed by a second number. The ? at the end means this part can appear 0 or 1 times. Inside it, (\d+) is the second capture group for the optional number.

Step 2: Python 2.7 Code Implementation

Since Python 2.7 doesn’t support f-strings, we’ll use str.format() instead to build the output. Here’s a complete working example:

import re

# Your example location strings
location_list = [
    "NC_000023.11:g.154532082",
    "NC_000023.11:g.154532058_154532060",
    "NC_000023.11:g.154532046"
]

# Precompile the regex for better performance with multiple strings
location_pattern = re.compile(r'NC_\d+\.\d+:g\.(\d+)(?:_(\d+))?')

for loc_str in location_list:
    match_result = location_pattern.match(loc_str)
    if match_result:
        first_pos = match_result.group(1)
        second_pos = match_result.group(2)
        
        # Generate the desired output format
        if second_pos:
            print("([{0}], [{1}])".format(first_pos, second_pos))
        else:
            print("([{0}])".format(first_pos))

Step 3: Expected Output

When you run this code, you’ll get exactly the output you requested:

([154532082])
([154532058], [154532060])
([154532046])

Quick Notes

  • This regex assumes all your input strings follow the exact format you provided. If you have edge cases (like different prefixes or more than two numbers), we can tweak the pattern easily.
  • Precompiling the regex with re.compile() is a good habit if you’re parsing a large number of strings—it speeds up the matching process.

内容的提问来源于stack exchange,提问作者Nulnulzeven

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:06:11