如何从Pandas DataFrame的坐标列中拆分提取x、y坐标值
问题原因
你遇到的TypeError: 'float' object is not iterable报错由列内的NaN缺失值直接导致:pandas默认将空值存储为float类型的np.nan,你之前用zip(*events['location'])拆分列时,遍历到float类型的空值就会因对象不可迭代抛出错误。
解决方案
通过统一的坐标拆分逻辑处理空值,同时支持批量处理所有同格式的坐标列,无需重复编写拆分代码:
- 先定义通用的坐标列拆分函数,自动识别空值并做兼容处理
import pandas as pd import numpy as np import json import glob import os def split_coord_col(df, col_name): # 将非列表类型的空值替换为[NaN, NaN],避免迭代报错,拆分后对应坐标自动为空值 temp_series = df[col_name].apply(lambda x: x if isinstance(x, list) else [np.nan, np.nan]) df[f"{col_name}.x"], df[f"{col_name}.y"] = zip(*temp_series) return df
- 保留你原有的JSON合并逻辑,合并完成后批量拆分所有需要处理的坐标列
path = os.fspath(r"C:\Users\user\Desktop\Football Data\Euro2020Matches") json_files = glob.glob(os.path.join(path, "*.json")) pd.set_option("display.max_columns", None) dfs = [] for file in json_files: with open(file, encoding='UTF-8') as f: json_data = pd.json_normalize(json.loads(f.read())) dfs.append(json_data) events = pd.concat(dfs, sort=False).reset_index(drop=True) # 把所有需要拆分的坐标列名放进列表即可批量处理 target_coord_cols = ["location", "pass.end_location", "carry.end_location"] for col in target_coord_cols: events = split_coord_col(events, col)
- 拆分完成后直接运行你原有的距离计算代码即可,空值对应的计算结果会自动返回NaN,不会影响其他有效行的运算
events['initialDistancefromgoal'] = np.sqrt(((120 - events['location.x'])**2) + ((40 -events['location.y'])**2)) events['finalDistancefromgoal'] = np.sqrt(((120 - events['pass.end_location.x'])**2) + ((40 - events['pass.end_location.y'])**2)) events['finalDistancefromgoalcarry'] = np.sqrt(((120 - events['carry.end_location.x'])**2) + ((40 - events['carry.end_location.y'])**2))
注:你之前写的循环读取JSON打印内容的代码属于调试用代码,不需要的话可以直接删除,不会影响最终的数据合并与计算逻辑。
内容的提问来源于stack exchange,提问作者Haythem Joudeh
相关产品推荐
相关产品推荐

