You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则提取指定首尾字符串及无序数据字段?解决正则报错

Fixing Your Python Regex Issues & Meeting Your Extraction Needs

Hey there! Let's tackle your problems step by step—first fixing that sre_constants.error: nothing to repeat error, then addressing both of your text extraction requirements.

Why You're Getting That Regex Error

That error almost always happens when your regex has a repeat quantifier (like *, +, ?, or {n,m}) with nothing to repeat in front of it. For example:

  • Accidentally writing *title: instead of title: (the * has no preceding character/group to act on)
  • Forgetting to escape special characters that are part of your field names (e.g., if a field was named page+dimensions, you'd need to write page\+dimensions instead of page+dimensions—the unescaped + is treated as a repeat quantifier here)

Double-check your regex strings for these mistakes, and that should clear up the error.

Your Extraction Requirements

1. Extract Strings from a List That Start/End With Specific Words

Suppose you have a list of options, and you want to pull out any string that starts with a word like Start and ends with End. Here's how to do it with regex:

import re

options = [
    "Start apple End",
    "Start banana End",
    "Orange Start cherry",
    "Start test End extra text"  # If you want to exclude extra text after End, adjust the regex
]

# Regex to match strings that START with "Start" and END with "End" (entire string)
pattern = re.compile(r'^Start.*?End$')
matches = [opt for opt in options if pattern.match(opt)]

# If you need to extract the content between Start and End (instead of the whole string)
content_matches = [re.search(r'Start(.*?)End', opt).group(1).strip() for opt in options if re.search(r'Start(.*?)End', opt)]

print("Full matches:", matches)
print("Content between Start/End:", content_matches)
  • ^ anchors the match to the start of the string
  • .*? is a non-greedy match for any characters in between
  • $ anchors the match to the end of the string

2. Extract Fields from Unordered, Messy Data

For your sample data (title: A Game of Thrones author: George R page dimensions: 210 x 297 mm) where fields can be in any order, you have two solid options:

Option 1: Extract Each Field Individually (Flexible for Any Order)

Use regex with positive lookahead to define where each field's value ends (either at the next field name or the end of the string):

import re

data = "author: George R page dimensions: 210 x 297 mm title: A Game of Thrones"  # Shuffled order

# Extract title
title = re.search(r'title:\s*(.*?)(?=\s+author:|\s+page dimensions:|$)', data).group(1)
# Extract author
author = re.search(r'author:\s*(.*?)(?=\s+title:|\s+page dimensions:|$)', data).group(1)
# Extract page dimensions
dimensions = re.search(r'page dimensions:\s*(.*?)(?=\s+title:|\s+author:|$)', data).group(1)

print(f"Title: {title}")
print(f"Author: {author}")
print(f"Page Dimensions: {dimensions}")

Option 2: Extract All Fields at Once (Cleaner for Multiple Fields)

Use a regex pattern to match all field-value pairs, then convert them to a dictionary for easy access:

import re

data = "page dimensions: 210 x 297 mm title: A Game of Thrones author: George R"

# Regex to match field names (including multi-word names like "page dimensions") and their values
pattern = re.compile(r'(\w+(?: \w+)?):\s*(.*?)(?=\s+\w+(?: \w+)?|:|$)')
field_dict = dict(pattern.findall(data))

print(field_dict)
# Output: {'page dimensions': '210 x 297 mm', 'title': 'A Game of Thrones', 'author': 'George R'}
  • (\w+(?: \w+)?) matches field names (supports single or multi-word names)
  • :\s* matches the colon and any leading whitespace after it
  • (.*?) captures the field value (non-greedy to stop at the next field)
  • (?=\s+\w+(?: \w+)?|:|$) is a positive lookahead that stops the value match when it hits the next field name or the end of the string

内容的提问来源于stack exchange,提问作者frorekable

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:31:22