如何拆分Pandas DataFrame中的双层嵌套列?解决索引越界报错
First, let's fix the "list index out of range" error you're encountering. The issue arises because some rows in your type column have empty lists—your lambda assumes every entry has at least one element at index 0, which isn't always true.
Fixing the Details Extraction Error
Modify your lambda to check if the list is non-empty before accessing its first element:
df['details'] = df['type'].apply(lambda x: x[0]['details'] if len(x) > 0 else None)
This returns None (converted to NaN in pandas) for rows with empty type lists, preventing the index error.
Handling the Nested Cause Field
Manual lambda extraction gets messy for deeply nested structures like Cause. Instead, use pandas' pd.json_normalize—it's designed to flatten nested JSON/dict structures into clean DataFrame columns.
Here's a step-by-step approach:
Convert
typelists to dicts:
Since eachtypeentry is a list (with 0 or 1 elements), convert it to a dict (empty dict if the list is empty):type_dicts = df['type'].apply(lambda x: x[0] if len(x) > 0 else {})Flatten the nested structure:
Usejson_normalizeto expand both the top-level fields intypeand the nestedCausefields:type_expanded = pd.json_normalize(type_dicts)This will create columns like
details,id,machine,Cause.code,Cause.description,Cause.id, andCause.reason.Merge with original DataFrame:
Combine the expandedtypedata with your original DataFrame, dropping the originaltypecolumn:final_df = pd.concat([df.drop('type', axis=1), type_expanded], axis=1)
Full Example with Your Sample Data
If we apply this to your sample entry, the final_df will have these columns:Speed, endTime, line, level, lineId, loss, startTime, details, id, machine, Cause.code, Cause.description, Cause.id, Cause.reason
This method is robust, handles missing fields automatically (filling with NaN), and scales well if your nested structure has more fields later on.
内容的提问来源于stack exchange,提问作者Keithx

