如何在Python中将DataFrame的geometry列拆分为多组x、y坐标列
DataFrame几何列拆分为多列的解决方案
你的代码出错有两个核心原因:一是没先去除geometry列里的polygon((和))冗余字符,直接拆分带包装的字符串会得到无效内容;二是拆分用的正则表达式写法错误。按以下步骤修正:
1. 清理冗余字符
先把geometry列中的空间类型标识和括号去掉,只保留纯坐标字符串:
df['clean_coords'] = df['geometry'].str[10:-2] # 从第10个字符开始截取,去掉最后2个字符,刚好剔除polygon((和))
如果字符串格式有变动,也可以用正则替换更灵活:
df['clean_coords'] = df['geometry'].str.replace(r'^polygon\(\(|\\)$', '', regex=True).str.strip()
2. 拆分坐标点
将清理后的字符串按, 拆分,得到独立的坐标点列:
point_cols = df['clean_coords'].str.split(', ', expand=True)
3. 拆分每个坐标的x、y值
遍历每个坐标点列,按空格拆分x和y,再合并成最终的多列结果:
# 初始化结果DataFrame final_result = pd.DataFrame() for col in point_cols.columns: # 拆分单个坐标的x和y xy_pair = point_cols[col].str.split(' ', expand=True) # 合并到结果中 final_result = pd.concat([final_result, xy_pair], axis=1) # 设置列名为重复的x、y final_result.columns = ['x', 'y'] * (final_result.shape[1] // 2)
完整可运行代码
import pandas as pd # 模拟你的原始DataFrame df = pd.DataFrame({ 'object_id': [1], 'shape': [450], 'geometry': ['polygon((6.6 51.2, 6.69 51.23, 6.69 51.2))'] }) # 执行清理与拆分流程 df['clean_coords'] = df['geometry'].str[10:-2] point_cols = df['clean_coords'].str.split(', ', expand=True) final_result = pd.DataFrame() for col in point_cols.columns: xy_pair = point_cols[col].str.split(' ', expand=True) final_result = pd.concat([final_result, xy_pair], axis=1) final_result.columns = ['x', 'y'] * (final_result.shape[1] // 2) print(final_result)
运行后输出完全符合你的需求:
x y x y x y 0 6.6 51.2 6.69 51.23 6.69 51.2
如果你的数据是空间数据(比如用GeoPandas存储),更规范的做法是直接提取几何对象的坐标:
import geopandas as gpd # 假设df是GeoDataFrame coords = df['geometry'].apply(lambda g: list(g.exterior.coords)) # 再将坐标展开成列即可
内容的提问来源于stack exchange,提问作者Gaurav Raina
相关产品推荐
相关产品推荐

