You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拆分LINESTRING坐标数据并生成指定格式的Pandas DataFrame?

Solution for Parsing LINESTRING Data in Pandas

Absolutely! Let's break down how to solve both of your requirements step by step:

Step 1: Fix the Initial Data Format

First, note that your original DataFrame example uses a raw LINESTRING keyword which isn't valid Python syntax—we'll assume your actual col1 values are string representations of the geometry (like "LINESTRING(...)"). Here's the corrected starting point:

import pandas as pd
import re

# Corrected input DataFrame (col1 contains string values)
d = {'col1': ["LINESTRING(174.76028 -36.80417,174.76041 -36.80389, 175.76232 -36.82345)"]}
df = pd.DataFrame(d)

Step 2: Define Your Custom Value-Processing Function

You can tailor this function to any logic you need (rounding, scaling, transformations, etc.). For example, let's use a function that rounds coordinates to 2 decimal places:

# Custom function to process individual coordinate values
def process_coord_value(val):
    # Example: Round to 2 decimal places; replace with your logic here
    return round(float(val), 2)

Step 3: Parse LINESTRING and Generate Target Columns

We'll create a helper function to extract the geometry type (like LINESTRING) and process each coordinate pair, then apply it to your DataFrame to split into col1 and col2:

def parse_linestring(geom_str):
    # Extract the geometry type (e.g., LINESTRING)
    geom_type = re.match(r'^(\w+)', geom_str).group(1)
    # Extract all coordinate pairs inside the parentheses
    coords_raw = re.findall(r'\((.*?)\)', geom_str)[0].split(',')
    
    # Process each coordinate pair and convert to tuples
    processed_coords = []
    for pair in coords_raw:
        x_str, y_str = pair.strip().split()
        processed_x = process_coord_value(x_str)
        processed_y = process_coord_value(y_str)
        processed_coords.append((processed_x, processed_y))
    
    return geom_type, processed_coords

# Apply the parser to col1 and expand into two new columns
df[['col1', 'col2']] = df['col1'].apply(lambda x: pd.Series(parse_linestring(x)))

Step 4: Verify the Result

Running the code above will give you exactly the format you requested:

print(df)
# Output:
#        col1                                                col2
# 0  LINESTRING  [(174.76, -36.8), (174.76, -36.8), (175.76, -36.82)]

Customization Tips

  • To change how values are processed, just modify the process_coord_value function (e.g., add 1 to each value, convert to integers, etc.).
  • If your data has other geometry types (like POLYGON), you can adjust the regex or add conditional logic to handle them.

内容的提问来源于stack exchange,提问作者NZ_DJ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:22:17