You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

从网络驱动器JSON提取步态参数时DataFrame为空的技术求助

问题描述

尝试从网络驱动器的JSON文件提取步态参数,运行Python代码后始终得到空DataFrame,需修改代码实现预期功能。

涉及函数说明

  • extract_data_from_json:读取JSON文件数据并返回统一格式
  • calculate_step_length_and_width:从字典(entry)中提取step_length和step_width
  • gaitsession_etl1:对步态指标数据执行ETL(抽取、转换、加载)操作

原代码

import os
import json
import pandas as pd
import numpy as np

def extract_data_from_json(filename):
    """Extract data from a JSON file."""
    try:
        with open(filename, 'r') as f:
            data = json.load(f)  # Load JSON data
            # Debug: Check if data is a dictionary or list
            if isinstance(data, dict):
                return [data]  # Wrap dictionary in a list for consistency
            elif isinstance(data, list):
                return data  # Return list directly if JSON data is a list
            else:
                print(f"Unexpected JSON data type in file {filename}: {type(data)}")
                return []
    except json.JSONDecodeError as e:
        print(f"Error decoding JSON file {filename}: {e}")
    except Exception as e:
        print(f"An unexpected error occurred with file {filename}: {e}")
    return []

def calculate_step_length_and_width(entry):
    """Calculate step length and width from entry data."""
    # Placeholder logic for calculation; replace with actual logic
    step_length = entry.get('step_length', np.nan)
    step_width = entry.get('step_width', np.nan)
    return step_length, step_width

def gaitsession_etl1():
    # Load session data
    df = pd.read_parquet('Z:\\dataframes\\df_session.parquet.gzip')

    # Select 'Laps Walking Application' sessions
    df_s = df[df['testName'].str.contains('Laps Walking Application')]

    if df_s.empty:
        print("No sessions found with 'Laps Walking Application'.")
        return

    patient_list = df_s['patient'].tolist()
    identifier_list = df_s['identifier'].tolist()

    results = []

    for i in range(len(patient_list)):
        path = f'Z:/raw_data/COVID Study FAU/patients/{patient_list[i]}/sessions/{identifier_list[i]}/files'
        name = f'{identifier_list[i]}_gaitmetrics.json'
        filename = os.path.join(path, name)
        
        # Debug prints
        print(f"Debug - Constructed path: {path}")
        print(f"Debug - Constructed filename: {filename}")
        
        if os.path.isfile(filename):
            data = extract_data_from_json(filename)
            
            if not data:
                print(f"No data found in file {filename}")
            
            for entry in data:
                if isinstance(entry, dict):
                    step_length, step_width = calculate_step_length_and_width(entry)
                    results.append({
                        'patient': patient_list[i],
                        'identifier': identifier_list[i],
                        'timestamp': entry.get('Timestamp', 'NaN'),
                        'step_length': step_length,
                        'step_width': step_width
                    })
                else:
                    print(f"Expected dictionary but got {type(entry)} in file {filename}")
        else:
            print(f"File not found: {filename}")
    
    data_gait = pd.DataFrame(results)

    if data_gait.empty:
        print("No data extracted. DataFrame is empty.")
    else:
        print(data_gait)

# Run the ETL function
gaitsession_etl1()

JSON文件示例

{
  "Steps": [
    {
      "AverageLeaning": "0",
      "stepLengthLeft": "0",
      "stepLengthRight": "0",
      "totalStepsLeft": "0",
      "totalStepsRight": "0",
      "totalSteps": "0",
      "averageSpeed": "0",
      "totalDistance": "0",
      "totalDistanceLeft": "0",
      "totalDistanceRight": "0",
      "swingStanceRatio": "0",
      "swingStanceRatioLeft": "0",
      "swingStanceRatioRight": "0",
      "averageCadence": "0",
      "averageCadenceL": "0",
      "averageCadenceR": "0",
      "averageStepLengthL": "0",
      "averageStepLengthR": "0",
      "averageStepLength": "0",
      "singleDoubleRatio": "0",
      "singleDoubleRatioLeft": "0",
      "singleDoubleRatioRight": "0",
      "FootSide": "1",
      "nombreFreezing": "0",
      "Arm": "0",
      "ArmAverage": "0",
      "Shoulder": "0",
      "ShoulderAverage": "0",
      "StepHeightL": "1700.4410231740912",
      "StepHeightR": "1161.1989661045372",
      "ShufflingPercentageL": "NaN",
      "ShufflingPercentageR": "NaN",
      "ShufflingPercentage": "NaN",
      "Timestamp": "00:01:52.31"
    }
  ]
}

修改方案

空DataFrame的核心原因是代码未正确解析JSON层级结构、使用了错误的参数键名,且未处理JSON中字符串类型的数值数据。以下是具体修改:

1. 修正extract_data_from_json函数,解析正确层级

JSON外层是字典,实际步态数据在Steps列表中,原函数直接返回外层字典的列表,无法获取有效数据。修改后:

def extract_data_from_json(filename):
    """Extract data from a JSON file."""
    try:
        with open(filename, 'r') as f:
            data = json.load(f)
            # 提取Steps列表,不存在则返回空列表
            if isinstance(data, dict) and 'Steps' in data:
                return data['Steps']
            elif isinstance(data, list):
                return data
            else:
                print(f"Unexpected JSON structure in file {filename}")
                return []
    except json.JSONDecodeError as e:
        print(f"Error decoding JSON file {filename}: {e}")
    except Exception as e:
        print(f"An unexpected error occurred with file {filename}: {e}")
    return []

2. 修正calculate_step_length_and_width函数,匹配正确键名并转换数据类型

JSON中的步长参数键名是averageStepLength/stepLengthLeft等,而非原代码的step_length;同时所有数值以字符串存储,需转换为数值类型(处理NaN字符串):

def calculate_step_length_and_width(entry):
    """Calculate step length and width from entry data."""
    # 提取平均步长,转换为数值
    avg_step_length_str = entry.get('averageStepLength', 'NaN')
    step_length = float(avg_step_length_str) if avg_step_length_str != 'NaN' else np.nan
    
    # 示例JSON无步宽数据,暂时设为NaN;若有相关键名(如stepWidthLeft),按上述逻辑提取
    step_width = np.nan
    
    return step_length, step_width

3. 优化循环逻辑(可选)

原代码通过索引遍历列表,改为同时遍历patient和identifier,更简洁:

# 替换原for循环部分
for patient, identifier in zip(patient_list, identifier_list):
    path = f'Z:/raw_data/COVID Study FAU/patients/{patient}/sessions/{identifier}/files'
    name = f'{identifier}_gaitmetrics.json'
    filename = os.path.join(path, name)
    
    print(f"Debug - Constructed filename: {filename}")
    
    if os.path.isfile(filename):
        data = extract_data_from_json(filename)
        
        if not data:
            print(f"No step data found in file {filename}")
            continue
        
        for entry in data:
            if isinstance(entry, dict):
                step_length, step_width = calculate_step_length_and_width(entry)
                results.append({
                    'patient': patient,
                    'identifier': identifier,
                    'timestamp': entry.get('Timestamp', 'NaN'),
                    'step_length': step_length,
                    'step_width': step_width
                })
            else:
                print(f"Expected dictionary entry in file {filename}, got {type(entry)}")
    else:
        print(f"File not found: {filename}")

4. 验证流程

修改后运行代码,检查debug输出:

  • 确认文件路径正确、文件存在
  • 确认extract_data_from_json返回了Steps列表中的数据
  • 确认calculate_step_length_and_width返回有效数值而非全部NaN

内容的提问来源于stack exchange,提问作者user26608122

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 02:15:55