You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Regular Expression提取名称及特定字符串上方的城镇数据?

Extracting Town Data Above a Specific String with Regex

Hey there! Let's walk through how to pull that town data you need, step by step. I'll use Python examples since it's super common for text processing, but the logic applies to most languages with regex support.

Step 1: Locate the Town Section Above Your Target String

First, you need to find where your target string lives in the text, then grab the town data right above it. The approach depends on how your text is structured:

Case 1: Text is split by separators (like --- or blank lines)

If your text has clear separators between entries, split the text into sections first, then find the section with your target string and grab the previous one:

# Sample text structure
text = """
Town: Springfield
Population: 15000
Region: North
---
Town: Rivertown
Population: 8000
Region: South
---
Order ID: 12345  # Your target string
"""

target_str = "Order ID: 12345"

# Split text into sections using your separator
sections = text.split("---")

# Find the index of the section containing your target string
target_index = next(i for i, sec in enumerate(sections) if target_str in sec)

# Grab the previous section (this has your town data)
town_section = sections[target_index - 1].strip()
print(town_section)
# Output: "Town: Springfield\nPopulation: 15000\nRegion: North"

Case 2: Text is line-separated (no clear section separators)

If your text is just a list of lines, find the line number of your target string, then scan upwards to find the line with "Town:":

lines = text.splitlines()
target_line_num = next(i for i, line in enumerate(lines) if target_str in line)

# Scan upwards from the target line to find the town entry
town_line = None
for i in range(target_line_num - 1, -1, -1):
    if "Town:" in lines[i]:
        town_line = lines[i].strip()
        break

if town_line:
    print(town_line)  # Output: "Town: Springfield"

Step 2: Extract the Town Name with Regular Expressions

Once you have the town section or line, use regex to pull out the name. The regex pattern will depend on what your town names look like (e.g., single word, multi-word, with hyphens):

Basic Regex for Town Names

This pattern works for most common town names (single or multi-word, no special characters):

import re

# Regex pattern: matches "Town: " followed by one or more words
town_pattern = r'Town: (\w+(?:\s\w+)*)'

# If you have the full town section
match = re.search(town_pattern, town_section)
# Or if you have just the town line: match = re.search(town_pattern, town_line)

if match:
    town_name = match.group(1)
    print(f"Extracted Town Name: {town_name}")  # Output: "Springfield"
else:
    print("No town name found in the section.")

Breaking Down the Regex

  • Town: : Matches the exact prefix before the town name (adjust this if your label is different, like City: or Village:)
  • (\w+(?:\s\w+)*): The capturing group that grabs the name:
    • \w+: Matches one or more word characters (letters, numbers, underscores)
    • (?:\s\w+)*: Matches zero or more instances of "space + word" (for multi-word names like "New Orleans")
    • ?: Makes this a non-capturing group, so we only get the full town name in group(1)

Adjustments for Special Characters

If your town names include hyphens or apostrophes (like "O'Neil" or "West-Haven"), update the regex to include those characters:

# Pattern for names with hyphens/apostrophes
town_pattern = r'Town: ([\w\'-]+(?:\s[\w\'-]+)*)'

Key Tips

  • Always test your regex against your actual text to make sure it covers all edge cases (like unusual town names)
  • If you're working in a different language (like JavaScript or Java), the regex logic stays the same—just adjust the code syntax for splitting text and running regex matches

内容的提问来源于stack exchange,提问作者Tom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:17:02