You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取带标签文本文件中的多边形图层、点数及坐标

Solution for Parsing Polygon Boundary Strings

Hey there! Let's work through this string parsing problem together—it looks like the core issue is reliably extracting the layer number, point count, and coordinate pairs from that specific boundary format, since basic whitespace splitting is picking up extra keywords and messing things up. Here's a step-by-step approach using regular expressions (which are perfect for this structured-but-uniformly-formatted text):

Step 1: Use Regular Expressions to Extract Key Components

First, we'll use a regex pattern to target exactly the values we need, ignoring the fixed keywords like layer, datatype, and xy. This avoids the problem of splitting on spaces and getting irrelevant text mixed in with our numbers.

Example Code

import re
import numpy as np
import pandas as pd

# Your sample input string
sample_boundary = "boundary layer 2 datatype 0 xy 15 525270 8663518 525280 8663520 525290 8663515 endel"

# Define regex pattern to capture layer number, point count, and coordinate string
pattern = r"boundary layer (\d+) datatype \d+ xy (\d+) (.*?) endel"
match = re.fullmatch(pattern, sample_boundary)

if match:
    # Extract and convert numeric values
    layer_number = int(match.group(1))
    point_count = int(match.group(2))
    coords_raw = match.group(3)
    
    # Split coordinate string into individual numbers and convert to integers
    coords_list = list(map(int, coords_raw.split()))
    
    # Convert to numpy array of (x, y) pairs
    coords_array = np.array(coords_list).reshape(-1, 2)
    
    # Create a pandas DataFrame (add layer number for context)
    df = pd.DataFrame(coords_array, columns=["x", "y"])
    df["layer"] = layer_number
    
    # Verify the results
    print(f"Layer Number: {layer_number}")
    print(f"Point Count: {point_count}")
    print("\nCoordinates DataFrame:")
    print(df)
else:
    print("The input string doesn't match the expected format.")

Why This Works Instead of Basic Splitting

  • The regex pattern (\d+) specifically captures numeric values for the layer and point count, skipping the fixed text around them.
  • The (.*?) captures all the coordinate values between xy [point_count] and endel (the non-greedy ? ensures we don't accidentally include extra text if there are unexpected characters).
  • Splitting the coordinate string after extraction works because we've already removed the irrelevant keywords—now we just have a list of x-y pairs separated by spaces.

Handling Multiple Boundary Entries

If you have a file or list with multiple boundary...endel blocks, you can loop through each entry and apply the same logic, appending the results to a single DataFrame or numpy array:

# Example list of boundary strings
boundary_list = [
    "boundary layer 2 datatype 0 xy 3 525270 8663518 525280 8663520 525290 8663515 endel",
    "boundary layer 3 datatype 0 xy 2 525300 8663500 525310 8663505 endel"
]

all_data = []
for entry in boundary_list:
    match = re.fullmatch(pattern, entry)
    if match:
        layer = int(match.group(1))
        coords = np.array(list(map(int, match.group(3).split()))).reshape(-1, 2)
        temp_df = pd.DataFrame(coords, columns=["x", "y"])
        temp_df["layer"] = layer
        all_data.append(temp_df)

# Combine all into one DataFrame
final_df = pd.concat(all_data, ignore_index=True)
print(final_df)

内容的提问来源于stack exchange,提问作者jax

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:23:38