You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用正则表达式拆分字符串并提取指定文件名前缀?

Solution for Extracting Prefix from Filenames Using Regex

Got it, let's break down how to solve this regex problem step by step. The key here is targeting the two fixed suffix patterns in your filenames and capturing everything before them.

The Regex Pattern

Here's the regex that will handle both filename formats:

^(.*?)(?:_P01)?-2I_114\.dexl\.gz$

Let's break down each part:

  • ^: Anchors the match to the start of the string, ensuring we don't pick up unwanted characters from the middle.
  • (.*?): This is a non-greedy capture group that grabs all characters until it hits the first occurrence of the suffix pattern. The ? makes it non-greedy, which is crucial here to avoid over-matching past the intended prefix.
  • (?:_P01)?: An optional non-capturing group that matches the _P01 segment if it exists. The ?: means we don't capture this part, and ? makes it optional.
  • -2I_114\.dexl\.gz$: Matches the fixed trailing part of the filename. We escape the dots with \ because dots in regex match any character by default. The $ anchors the match to the end of the string to ensure we're matching the full filename.

Example Usage

Let's test this pattern against your sample filenames:

  • Input: gb_reg_test2-2I_114.dexl.gz → Captured prefix: gb_reg_test2
  • Input: gb_bk_test1_P01-2I_114.dexl.gz → Captured prefix: gb_bk_test1
  • Input: aa_bb_cc-2I_114.dexl.gz → Captured prefix: aa_bb_cc

Code Example (Python)

If you're using Python to implement this, here's a quick snippet:

import re

# Define the regex pattern
filename_pattern = r'^(.*?)(?:_P01)?-2I_114\.dexl\.gz$'

# List of sample filenames
sample_filenames = [
    'gb_reg_test2-2I_114.dexl.gz',
    'gb_bk_test1_P01-2I_114.dexl.gz',
    'aa_bb_cc-2I_114.dexl.gz'
]

# Extract prefixes
for name in sample_filenames:
    match_result = re.match(filename_pattern, name)
    if match_result:
        print(f"Filename: {name} → Extracted prefix: {match_result.group(1)}")

Running this code will output:

Filename: gb_reg_test2-2I_114.dexl.gz → Extracted prefix: gb_reg_test2
Filename: gb_bk_test1_P01-2I_114.dexl.gz → Extracted prefix: gb_bk_test1
Filename: aa_bb_cc-2I_114.dexl.gz → Extracted prefix: aa_bb_cc

This regex works because it's specifically tailored to the two suffix variants you have, and the non-greedy match ensures we stop exactly at the start of the fixed suffix, regardless of how many underscores are in your prefix.

内容的提问来源于stack exchange,提问作者GBMan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:29:18