将JSON文件内容提取为DataFrame列时遇错误求助
Hey there! Let's tackle those frustrating errors you're hitting when trying to extract specific parts of a JSON file into a DataFrame. Those row mismatch and type coercion issues are super common with nested JSON structures, so let's break down what's going wrong and how to fix it.
Why You're Seeing Those Errors
First, let's unpack the two main issues:
arguments imply differing number of rows: 0, 1: This happens because when you usec()to pull nested JSON fields likevendor_dataorproblemtype_data, you're flattening lists into vectors—but some CVE entries might be missing one of these fields, leading to mismatched lengths between yourdf,dfone, anddftwoobjects.cannot coerce type 'closure' to vector of type 'list': Most likely, you accidentally reused a variable name that's already an R built-in function (likedf—if you ever assigneddf <- data.frameby mistake, R will treatdfas the function instead of your data object).
Solutions to Extract & Merge Your Data Correctly
Instead of using c() to flatten lists, use tools designed for nested data structures like purrr or jsonlite to ensure consistent row counts and proper data types.
Approach 1: Use purrr to Map Nested Lists to DataFrames
This method lets you explicitly handle each nested field and ensure every CVE entry gets a row (even if some fields are missing, they'll fill with NA):
# Load required packages library(purrr) library(dplyr) # Extract vendor data: convert each nested vendor_data list to a row in a DataFrame vendor_df <- map_dfr(x$CVE_Items$cve$affects$vendor$vendor_data, ~as.data.frame(.x, stringsAsFactors = FALSE)) # Extract problem type data problemtype_df <- map_dfr(x$CVE_Items$cve$problemtype$problemtype_data, ~as.data.frame(.x, stringsAsFactors = FALSE)) # Extract description data (filling in the missing part of your code) description_df <- map_dfr(x$CVE_Items$cve$description$description_data, ~as.data.frame(.x, stringsAsFactors = FALSE)) # Merge all DataFrames (they'll have matching row counts now!) final_df <- cbind(vendor_df, problemtype_df, description_df)
Approach 2: Let jsonlite Flatten the JSON Automatically
If your JSON structure is consistent, jsonlite can do the heavy lifting by flattening nested fields into easy-to-select columns:
library(jsonlite) # Read the JSON file and flatten nested structures into dot-separated column names x <- fromJSON("your_cve_file.json", flatten = TRUE) # Select only the columns you need from the flattened CVE_Items data final_df <- x$CVE_Items %>% select( starts_with("cve.affects.vendor.vendor_data"), starts_with("cve.problemtype.problemtype_data"), starts_with("cve.description.description_data") )
Quick Tips to Avoid Future Errors
- Avoid generic variable names: Instead of
df, use specific names likevendor_listorproblemtype_dfto prevent conflicts with R's built-in functions. - Check nested structure first: Use
str(x$CVE_Items)to inspect the JSON hierarchy—this helps you spot which fields might be missing in some entries. - Use
map_dfrfor list-to-DataFrame conversion: It automatically handles missing entries and ensures all outputs have the same number of rows.
内容的提问来源于stack exchange,提问作者user9664399

