R语言合并两个JSON时间序列数据问题求助
Hey there! Let's break down what's going wrong and fix your problem step by step—since you're new to R, you ran into a few common pitfalls with data structures and merging, which we can sort out easily.
First: Let's Fix the Core Merging Problem
Your goal is a full join: combine both datasets so every period from either A or B is included, with NA where a value is missing from one dataset. Here's the correct approach, starting with properly extracting your data:
Step 1: Load Packages & Extract Data Correctly
When you use jsonlite::fromJSON() on those API links, the $data$values part is already a ready-to-use data frame (exactly the table structure you need!). No need to mess with rbind() or t()—that's where your code went off track.
# Load required packages library(jsonlite) library(dplyr) # For easy renaming and merging (optional but cleaner) # Get Data A and rename the value column to avoid confusion data_final_A <- fromJSON("https://api.db.nomics.world/api/v1/json/series/imf-weo-ngap-npgdp-fra-6") df_A <- data_final_A$data$values %>% rename(value_A = value) # Rename "value" to "value_A" # Get Data B and do the same data_final_B <- fromJSON("https://api.db.nomics.world/api/v1/json/series/oecd-eo-fra-gap-a") df_B <- data_final_B$data$values %>% rename(value_B = value) # Rename "value" to "value_B"
Step 2: Merge the Data Frames
Use full_join() (from dplyr) to get all periods from both datasets, with NA for missing values. If you prefer base R, use merge() with all = TRUE:
# Option 1: Using dplyr (cleaner syntax) merged_df <- full_join(df_A, df_B, by = "period") # Option 2: Using base R (no extra package needed) merged_df <- merge(df_A, df_B, by = "period", all = TRUE)
If you run head(merged_df) or tail(merged_df) now, you'll see exactly what you want: all periods, with value_A and value_B filled where available, and NA where not.
Why Your Original Code Failed
Let's unpack the mistakes in your original code to help you avoid them later:
- Incorrect data extraction: You tried to build a table manually with
rbind()andt(), then messed up row names. The API already gives you a data frame—use that directly! - Typo in value extraction: You wrote
value_A<-data_final_B$data$values$value(pulling from B instead of A) which would have mixed up your data even before merging. - Missing
byparameter inmerge(): When you don't specifyby = "period", R tries to merge on all columns with matching names. Since your messed-up tables had generic column names likeV1andV2, it created a giant Cartesian product (every row from A paired with every row from B) instead of matching by period.
Clarifying List, Data Frame, and Table
You mentioned confusion here—let's simplify:
- List: A flexible container that can hold any type of data (vectors, data frames, other lists). The object you get from
fromJSON()is a list, which contains nested elements likedataandvalues. - Data Frame: A special type of list where every element (column) is a vector of the same length. Think of it like an Excel spreadsheet—this is the structure you should use for your time series data (rows = observations, columns = variables like
periodandvalue). - Table: A structure specifically for counting frequencies (e.g., how many times each value appears in a vector). It's not meant for storing raw time series data, so using names like
table_Awas a bit misleading here.
Final Check
After running the corrected code, your merged data frame will:
- Include every
periodfrom both datasets - Have
value_Afilled for periods only in Data A,value_Bfilled for periods only in Data B - Show both values for overlapping periods
- Use
NAfor missing values exactly as you wanted
内容的提问来源于stack exchange,提问作者Tochoka

