R语言For循环创建变量报错及合并数据帧问题求助
Hey there! Let's work through your problem step by step—first fixing that error, then making your code cleaner and more efficient.
First: Why You're Getting That Error
You hit the nail on the head: nam is just a string (like "cik_0000320193"), not the actual data frame you created with assign(). When you try nam[i,] <- ..., R doesn't recognize it as a data frame—it thinks you're trying to subscript a string, which doesn't make sense.
A quick fix for your existing loop would be to explicitly retrieve the data frame, modify it, then save it back:
for (j in cik_sample){ c_filings = company_filings(j, type = "10-K",count = 1000 ) row_num = nrow(c_filings) if (row_num == 0){ next } c_filings_short = select(c_filings, filing_date, href, type) c_hrefs <- as.data.frame(c_filings_short[,2]) c_length = length(c_hrefs) c_index = (1:c_length) nam <- paste("cik_", j,sep="") assign(nam, data.frame(matrix(ncol =21, nrow = c_length ))) # Fix the inner loop here for (i in c_index) { # Get the actual data frame from the global environment current_df <- get(nam) # Fill the ith row current_df[i,] <- filing_filers(as.character(c_hrefs[i,1])) # Save the updated data frame back assign(nam, current_df) } }
But honestly, using assign() and managing separate data frames for each CIK is messy and error-prone. Let's look at a better approach.
Better Approach: Use Lists + Tidyverse Tools
Instead of creating separate data frames, we can store results in a list and then combine everything into one big data frame automatically. This uses purrr for iterating and dplyr for data manipulation—way cleaner!
First, load the required packages:
library(EDGARWebR) library(dplyr) library(purrr) library(lubridate) # For extracting years from dates
Then write a function to process a single CIK:
process_single_cik <- function(cik) { # Get all 10-K filings for the CIK filings <- company_filings(cik, type = "10-K", count = 1000) # Skip if no filings exist if (nrow(filings) == 0) { return(tibble()) # Return empty tibble to keep things consistent } # Keep only the columns we need, add CIK and year from filing date filings_clean <- filings %>% select(filing_date, href) %>% mutate( cik = cik, filing_year = year(filing_date) ) # Safely fetch filer data (handle cases where a link fails) safe_get_filers <- possibly(filing_filers, otherwise = tibble()) # Iterate over each href, get filer data, and combine with filing info filings_clean %>% mutate(filer_data = map(href, safe_get_filers)) %>% unnest(filer_data, keep_empty = TRUE) # Keep rows even if filer data is missing }
Now process all your CIKs and combine into one data frame:
cik_sample <- c("0001438823","0000320193","0000018230") # Process every CIK and row-bind the results final_combined_df <- map_dfr(cik_sample, process_single_cik)
Why This Is Better
- No more
assign()or messy global environment variables—all results are managed in lists/data frames. map_dfr()automatically combines all results into a single data frame, so you don't have to manually merge later.possibly()adds error handling: if a singlefiling_filers()call fails (e.g., broken link), it won't crash the entire process.- We added a
filing_yearcolumn to make it easier to track which year each filing corresponds to.
Merging Your Existing Data Frames (If You Still Want To)
If you already ran your original loop and have those cik_xxxx data frames, you can combine them like this:
# Get all data frame names starting with "cik_" cik_df_names <- ls(pattern = "^cik_") # Combine them into one data frame combined_df <- bind_rows(mget(cik_df_names))
内容的提问来源于stack exchange,提问作者jayjunior

