R语言不等长嵌套列表转数据框报错问题求助
Hey there, let's work through this stock data conversion problem together—it’s super common to hit snags with nested lists and wonky download glitches, so you’re not alone!
First, let's zero in on that wonky "v..." volume entry from 2018-01-01. This is almost certainly a download truncation or formatting error.
First, check the problematic entry directly to confirm:
# View the 2018-01-01 entry data[["2018-01-01"]]
Then, let's clean up this field across your entire list. We'll replace any "v..." values with NA (or you can manually input the correct volume if you have it) and ensure the volume is treated as a numeric value:
library(purrr) # Clean up the volume field in every list entry data_fixed <- map(data, function(stock_entry) { # Replace "v..." with NA, then convert to numeric stock_entry$volume <- ifelse(grepl("v\\.\\.\\.", stock_entry$volume), NA_real_, as.numeric(stock_entry$volume)) stock_entry })
If you want to re-download just that date's data instead of using NA, packages like tidyquant make this easy:
library(tidyquant) # Re-fetch data for the problematic date fixed_date_data <- tq_get("YOUR_STOCK_SYMBOL", from = "2018-01-01", to = "2018-01-01") # Replace the bad entry in your list data[["2018-01-01"]] <- as.list(fixed_date_data)
Now that the bad data is fixed, let's turn that nested list into a clean data frame. The purrr::map_dfr function is perfect for this—it will stack each list entry as a row, and we can add a date column using the list names:
library(dplyr) # Convert list to data frame, with date as a column stock_df <- map_dfr(data_fixed, ~as.data.frame(.x), .id = "date")
If some of your list entries have more/fewer fields than others (a common issue with scraped data), we can standardize the columns first to avoid conversion errors:
# Get all unique column names across the entire list all_columns <- unique(unlist(map(data_fixed, names))) # Make every list entry have the same columns (fill missing ones with NA) data_uniform <- map(data_fixed, function(entry) { # Add missing columns with NA missing_cols <- setdiff(all_columns, names(entry)) entry[missing_cols] <- NA # Reorder columns to match the full set entry[all_columns] }) # Now convert to data frame stock_df <- map_dfr(data_uniform, ~as.data.frame(.x), .id = "date")
Once your data frame is clean, importing to MySQL is straightforward with the DBI and RMySQL packages:
library(DBI) library(RMySQL) # Connect to your MySQL database db_conn <- dbConnect(MySQL(), dbname = "your_database_name", host = "your_host_address", user = "your_username", password = "your_password") # Write the data frame to a table (overwrite if it already exists) dbWriteTable(db_conn, name = "stock_data", value = stock_df, overwrite = TRUE) # Don't forget to close the connection! dbDisconnect(db_conn)
A quick sanity check before importing: run str(stock_df) to make sure all numeric fields (like volume, open, close) are indeed numeric types—this will prevent weird errors in MySQL.
内容的提问来源于stack exchange,提问作者bchan

