You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python 2.7中根据指定标签提取字符串中的对应数据

Solution for Extracting Labeled Data in Python 2.7

Hey there! For extracting values by their labels in Python 2.7 from your formatted string, regex is a solid approach here—especially since some labels include spaces (like RETURN DATE). Here's a straightforward, reusable way to do it:

Step 1: Define the Source String and Extraction Function

First, we'll use Python's re module to create a function that targets specific labels and pulls their associated values. The regex pattern is designed to handle labels with spaces and stop at the next labeled field or the end of the string.

import re

# Your input string
source_str = "DATE: 7/25/2017 DATE OPENED: 7/25/2017 RETURN DATE: 7/26/2017 NUMBER: 201707250008754 RATE: 10.00"

def get_value_by_label(input_str, target_label):
    # Escape special regex characters in the label (in case of future special chars)
    escaped_label = re.escape(target_label)
    # Regex pattern: match label, colon, whitespace, then capture content until next label or end
    pattern = r'{0}:\s*(.*?)(?=\s+[A-Z]|$)'.format(escaped_label)
    match = re.search(pattern, input_str)
    # Return the matched value or None if label not found
    return match.group(1).strip() if match else None

Step 2: Test the Function

You can call this function with any of your target labels to get the corresponding value:

# Test cases
print(get_value_by_label(source_str, "DATE"))  # Output: 7/25/2017
print(get_value_by_label(source_str, "RETURN DATE"))  # Output: 7/26/2017
print(get_value_by_label(source_str, "NUMBER"))  # Output: 201707250008754
print(get_value_by_label(source_str, "RATE"))  # Output: 10.00
print(get_value_by_label(source_str, "DATE OPENED"))  # Output: 7/25/2017

How the Regex Works

Let's break down the pattern to make it clear:

  • {0}:\s*: Matches the target label followed by a colon and any amount of whitespace.
  • (.*?): Non-greedy capture group that grabs everything until the next stopping point (prevents over-matching).
  • (?=\s+[A-Z]|$): Positive lookahead that stops the capture when it encounters:
    • A space followed by an uppercase letter (the start of the next label), or
    • The end of the string (for the last field like RATE).

Edge Cases Handled

  • Labels with spaces (e.g., RETURN DATE, DATE OPENED)
  • Fields at the end of the string (like RATE: 10.00)
  • Extra whitespace around values (the .strip() ensures clean output)

内容的提问来源于stack exchange,提问作者Ravi K

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 09:03:29