You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何创建含变量与注释的正则表达式?现有实现失效求助

Fixing Your Regex with Comments and Variables

Let's break down why your original regex wasn't working and fix it step by step.

Key Issues in Your Original Code

  1. Comments were swallowing your pattern: When you concatenated the strings without line breaks, the # study number comment continued through your doctype variable and into the next comment. In re.VERBOSE mode, anything after a # (not inside a character class) is treated as a comment until the end of the line—so your doctype and version sections were never actually part of the regex pattern!
  2. Missing literal separators: VERBOSE mode ignores whitespace, so you need to explicitly include any required characters (like hyphens) between sections of your target string.
  3. Unescaped variable risk: If your doctype contains regex special characters (e.g., ., *, +), it would break the regex unless escaped.

Corrected Code

import re

# Example value for your doctype variable
doctype = "report"

# Properly structured regex with comments, variables, and VERBOSE mode
regex = re.compile(
    r'''
    ^(sm\d{5}-[a-z]+-\d{2})  # Capture the study number (format: smXXXXX-XXXX-XX)
    -                         # Hyphen separator between study number and doctype
    ''' + re.escape(doctype) + r'''  # Insert the document type variable (escaped for safety)
    -                         # Hyphen separator between doctype and version
    v(\d+)-(\d+)              # Capture major and minor version numbers (format: vXX-XX)
    \.pdf$                    # Match the literal .pdf file extension
    ''',
    re.VERBOSE
)

# Example usage
test_string = "sm12345-test-01-report-v01-02.pdf"
match = regex.match(test_string)

if match:
    print(f"Study Number: {match.group(1)}")
    print(f"Major Version: {match.group(2)}")
    print(f"Minor Version: {match.group(3)}")

What Changed?

  • Line breaks for comments: Each comment is now on its own line, so they don't interfere with the actual pattern parts.
  • Explicit separators: Added - between sections (adjust this if your target string uses a different separator like _ or nothing).
  • Escaped variable: Used re.escape(doctype) to ensure any special characters in your variable are treated as literal text, not regex operators.
  • Cleaner structure: The regex is now easy to read and modify, thanks to VERBOSE mode and well-placed comments.

If your target string doesn't use hyphens between sections, just remove or replace the - characters in the pattern to match your actual format.

内容的提问来源于stack exchange,提问作者Manar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:55:04