You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用正则表达式提取字符串中的description字段内容?

Hey there! Let's figure out how to extract all description field values from your string containing nested JSON, while ensuring all code snippets, technical terms, and special formatting stay exactly as they are. Here's a practical, step-by-step solution:

Solution for Extracting description Fields with Preserved Formatting

1. Craft a Recursive Regex Pattern

Since your string has nested JSON, we need a regex that can handle nested objects/arrays (not just flat strings). This pattern will match the description field and capture its value, regardless of whether it's a quoted string, nested object, or array:

/"description"\s*:\s*(?:"(?:\\.|[^"\\])*"|\{(?:[^{}]|(?R))*\}|\[(?:[^\[\]]|(?R))*\])/g

Breakdown of the Pattern:

  • "description"\s*:\s*: Matches the field name, plus any whitespace around the colon.
  • The capture group handles three scenarios:
    • "([^"\\]|\\.)*": Catches quoted strings, including those with escaped quotes (like \").
    • \{(?:[^{}]|(?R))*\}: Recursively matches nested JSON objects.
    • \[(?:[^\[\]]|(?R))*\]: Recursively matches nested JSON arrays.

2. Example Implementation (Python)

Python's standard re module doesn't support recursive regex, so we'll use the regex library (install it first with pip install regex). This code extracts all description values and cleans up quoted strings while preserving formatting:

import regex

# Replace this with your actual input string
input_str = """
{
  "item": 1,
  "description": "Supports `code blocks`, *emphasized text*, and special chars like !@#$",
  "nested_data": {
    "sub_item": 2,
    "description": "Nested description with [custom links] and nested JSON: {\"key\": \"value\"}"
  }
}
"""

# Our regex pattern
pattern = r'"description"\s*:\s*(?:"(?:\\.|[^"\\])*"|\{(?:[^{}]|(?R))*\}|\[(?:[^\[\]]|(?R))*\])'
matches = regex.findall(pattern, input_str)

# Process matches to clean up quoted strings (keep nested structures intact)
extracted_descriptions = []
for match in matches:
    if match.startswith('"') and match.endswith('"'):
        # Remove surrounding quotes and unescape escaped characters
        cleaned = match[1:-1].replace('\\"', '"').replace('\\\\', '\\')
        extracted_descriptions.append(cleaned)
    else:
        # Keep nested objects/arrays as-is
        extracted_descriptions.append(match)

# Output the results
for idx, desc in enumerate(extracted_descriptions, 1):
    print(f"Description {idx}: {desc}")

3. Critical Tips to Preserve Formatting

  • Avoid over-processing: Only unescape quotes for string values — don't modify code ticks, asterisks, links, or special characters.
  • Alternative for well-formed JSON: If your input is valid JSON, parsing it fully (using json.loads in Python or JSON.parse in JavaScript) and traversing the structure to collect description fields is more reliable for deeply nested data. Regex is great for quick scraping when you don't need full JSON validation.
  • Cross-language support: This recursive regex works in JavaScript (native RegExp), PHP, and other languages that support recursive regex syntax.

内容的提问来源于stack exchange,提问作者Abhishek Sharma M

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:43:05