从网络驱动器JSON提取步态参数时DataFrame为空的技术求助
问题描述
尝试从网络驱动器的JSON文件提取步态参数,运行Python代码后始终得到空DataFrame,需修改代码实现预期功能。
涉及函数说明
extract_data_from_json:读取JSON文件数据并返回统一格式calculate_step_length_and_width:从字典(entry)中提取step_length和step_widthgaitsession_etl1:对步态指标数据执行ETL(抽取、转换、加载)操作
原代码
import os import json import pandas as pd import numpy as np def extract_data_from_json(filename): """Extract data from a JSON file.""" try: with open(filename, 'r') as f: data = json.load(f) # Load JSON data # Debug: Check if data is a dictionary or list if isinstance(data, dict): return [data] # Wrap dictionary in a list for consistency elif isinstance(data, list): return data # Return list directly if JSON data is a list else: print(f"Unexpected JSON data type in file {filename}: {type(data)}") return [] except json.JSONDecodeError as e: print(f"Error decoding JSON file {filename}: {e}") except Exception as e: print(f"An unexpected error occurred with file {filename}: {e}") return [] def calculate_step_length_and_width(entry): """Calculate step length and width from entry data.""" # Placeholder logic for calculation; replace with actual logic step_length = entry.get('step_length', np.nan) step_width = entry.get('step_width', np.nan) return step_length, step_width def gaitsession_etl1(): # Load session data df = pd.read_parquet('Z:\\dataframes\\df_session.parquet.gzip') # Select 'Laps Walking Application' sessions df_s = df[df['testName'].str.contains('Laps Walking Application')] if df_s.empty: print("No sessions found with 'Laps Walking Application'.") return patient_list = df_s['patient'].tolist() identifier_list = df_s['identifier'].tolist() results = [] for i in range(len(patient_list)): path = f'Z:/raw_data/COVID Study FAU/patients/{patient_list[i]}/sessions/{identifier_list[i]}/files' name = f'{identifier_list[i]}_gaitmetrics.json' filename = os.path.join(path, name) # Debug prints print(f"Debug - Constructed path: {path}") print(f"Debug - Constructed filename: {filename}") if os.path.isfile(filename): data = extract_data_from_json(filename) if not data: print(f"No data found in file {filename}") for entry in data: if isinstance(entry, dict): step_length, step_width = calculate_step_length_and_width(entry) results.append({ 'patient': patient_list[i], 'identifier': identifier_list[i], 'timestamp': entry.get('Timestamp', 'NaN'), 'step_length': step_length, 'step_width': step_width }) else: print(f"Expected dictionary but got {type(entry)} in file {filename}") else: print(f"File not found: {filename}") data_gait = pd.DataFrame(results) if data_gait.empty: print("No data extracted. DataFrame is empty.") else: print(data_gait) # Run the ETL function gaitsession_etl1()
JSON文件示例
{ "Steps": [ { "AverageLeaning": "0", "stepLengthLeft": "0", "stepLengthRight": "0", "totalStepsLeft": "0", "totalStepsRight": "0", "totalSteps": "0", "averageSpeed": "0", "totalDistance": "0", "totalDistanceLeft": "0", "totalDistanceRight": "0", "swingStanceRatio": "0", "swingStanceRatioLeft": "0", "swingStanceRatioRight": "0", "averageCadence": "0", "averageCadenceL": "0", "averageCadenceR": "0", "averageStepLengthL": "0", "averageStepLengthR": "0", "averageStepLength": "0", "singleDoubleRatio": "0", "singleDoubleRatioLeft": "0", "singleDoubleRatioRight": "0", "FootSide": "1", "nombreFreezing": "0", "Arm": "0", "ArmAverage": "0", "Shoulder": "0", "ShoulderAverage": "0", "StepHeightL": "1700.4410231740912", "StepHeightR": "1161.1989661045372", "ShufflingPercentageL": "NaN", "ShufflingPercentageR": "NaN", "ShufflingPercentage": "NaN", "Timestamp": "00:01:52.31" } ] }
修改方案
空DataFrame的核心原因是代码未正确解析JSON层级结构、使用了错误的参数键名,且未处理JSON中字符串类型的数值数据。以下是具体修改:
1. 修正extract_data_from_json函数,解析正确层级
JSON外层是字典,实际步态数据在Steps列表中,原函数直接返回外层字典的列表,无法获取有效数据。修改后:
def extract_data_from_json(filename): """Extract data from a JSON file.""" try: with open(filename, 'r') as f: data = json.load(f) # 提取Steps列表,不存在则返回空列表 if isinstance(data, dict) and 'Steps' in data: return data['Steps'] elif isinstance(data, list): return data else: print(f"Unexpected JSON structure in file {filename}") return [] except json.JSONDecodeError as e: print(f"Error decoding JSON file {filename}: {e}") except Exception as e: print(f"An unexpected error occurred with file {filename}: {e}") return []
2. 修正calculate_step_length_and_width函数,匹配正确键名并转换数据类型
JSON中的步长参数键名是averageStepLength/stepLengthLeft等,而非原代码的step_length;同时所有数值以字符串存储,需转换为数值类型(处理NaN字符串):
def calculate_step_length_and_width(entry): """Calculate step length and width from entry data.""" # 提取平均步长,转换为数值 avg_step_length_str = entry.get('averageStepLength', 'NaN') step_length = float(avg_step_length_str) if avg_step_length_str != 'NaN' else np.nan # 示例JSON无步宽数据,暂时设为NaN;若有相关键名(如stepWidthLeft),按上述逻辑提取 step_width = np.nan return step_length, step_width
3. 优化循环逻辑(可选)
原代码通过索引遍历列表,改为同时遍历patient和identifier,更简洁:
# 替换原for循环部分 for patient, identifier in zip(patient_list, identifier_list): path = f'Z:/raw_data/COVID Study FAU/patients/{patient}/sessions/{identifier}/files' name = f'{identifier}_gaitmetrics.json' filename = os.path.join(path, name) print(f"Debug - Constructed filename: {filename}") if os.path.isfile(filename): data = extract_data_from_json(filename) if not data: print(f"No step data found in file {filename}") continue for entry in data: if isinstance(entry, dict): step_length, step_width = calculate_step_length_and_width(entry) results.append({ 'patient': patient, 'identifier': identifier, 'timestamp': entry.get('Timestamp', 'NaN'), 'step_length': step_length, 'step_width': step_width }) else: print(f"Expected dictionary entry in file {filename}, got {type(entry)}") else: print(f"File not found: {filename}")
4. 验证流程
修改后运行代码,检查debug输出:
- 确认文件路径正确、文件存在
- 确认
extract_data_from_json返回了Steps列表中的数据 - 确认
calculate_step_length_and_width返回有效数值而非全部NaN
内容的提问来源于stack exchange,提问作者user26608122
相关产品推荐
相关产品推荐

