机器学习项目中JSON转CSV:拆分多值pos特征为pos_x、pos_y、pos_z
Hey there! I’ve run into this exact scenario before—dealing with JSON arrays that turn into stringified lists in CSV is super common when working with ML datasets. Let’s get your pos feature split into pos_x, pos_y, and pos_z properly.
Option 1: Split During JSON → CSV Conversion (Best Practice)
Instead of letting the pos array turn into a string in the CSV, we can split it directly while building the DataFrame from your JSON data. This saves you from having to clean up the CSV later.
Here’s how to modify your existing code:
import pandas as pd import json data = [] with open('JSONfile.json') as fh: for line in fh: data.append(json.loads(line)) df = pd.DataFrame.from_dict(data) # Split the 'pos' list into three separate columns pos_columns = pd.DataFrame(df['pos'].tolist(), columns=['pos_x', 'pos_y', 'pos_z']) # Merge the new columns with the original DataFrame, dropping the old 'pos' column df = pd.concat([df.drop('pos', axis=1), pos_columns], axis=1) # Save the cleaned CSV df.to_csv('csvFile.csv', index=False)
Option 2: Fix an Existing CSV with Stringified 'pos' Arrays
If you already have the CSV where pos is stored as a string (like "[3838.387..., 5853.151..., 1.895]"), we can parse those strings back into lists and split them:
import pandas as pd import ast # Safe alternative to eval() for parsing literals # Load the existing CSV df = pd.read_csv('csvFile.csv') # Convert the stringified lists back into actual Python lists # ast.literal_eval is safer than eval() because it only parses valid Python literals df['pos'] = df['pos'].apply(ast.literal_eval) # Split into separate columns just like before pos_columns = pd.DataFrame(df['pos'].tolist(), columns=['pos_x', 'pos_y', 'pos_z']) df = pd.concat([df.drop('pos', axis=1), pos_columns], axis=1) # Save the processed CSV df.to_csv('processed_csvFile.csv', index=False)
Handling Missing Values (Optional)
If your dataset has rows where pos might be missing or invalid, you can add a safety check to avoid errors:
def parse_pos(pos_str): try: return ast.literal_eval(pos_str) if pd.notna(pos_str) else [None, None, None] except (ValueError, SyntaxError): return [None, None, None] df['pos'] = df['pos'].apply(parse_pos)
This will replace any invalid or missing pos entries with None values, which are easier to handle in ML workflows.
内容的提问来源于stack exchange,提问作者selima

