解决R语言爬虫中length(url) == 1 is not TRUE错误
批量爬取电网数据时URL拼接引发的
length(url) == 1 is not TRUE错误修复 我用R语言爬取smartgriddashboard.com的发电数据,通过开发者工具找到单日数据接口:https://www.smartgriddashboard.com/DashboardService.svc/data?area=generationactual®ion=ALL&datefrom=28-Sep-2023+00%3A00&dateto=28-Sep-2023+23%3A59
原本想循环遍历2022、2023年的所有月份批量拉取整月数据,但运行代码时触发错误:
length(url) == 1 is not TRUE
排查后确定是URL拼接环节出了问题,以下是原代码:
# Initialise empty list and vector tempds <- list() months_years <- character(0) years <- c('2022', '2023') months <- c('Jan','Feb', 'Mar', 'Apr','May','June','July','Aug','Sept','Oct','Nov','Dec') last_day <- c('31','28','31','30','31','30','31','31','30','31','30','31') # loop through years and months for (year in years){ i = 0 base_url <- "https://www.smartgriddashboard.com/DashboardService.svc/data?area=generationactual®ion=ALL&datefrom=01-', month, '-', year,+00%3A00&dateto='" url_end <- paste0('%2023:59') for(month in months) { print(paste(month, year)) url <- paste0(base_url, last_day, '-', month, '-', year, url_end) cat(url, " ") print(url) if (year == '2023' && month == 'Sept'){ print(paste(month, year, 'has been reached')) break } # Send GET request and read CSV data response <- GET(url) if (http_status(response)$status_code == 200) { content <- content(response, "text") mds <- read_csv(content, na = "-") # Process the data mds$date <- as.POSIXct(mds$`DATE & TIME`, format="%d-%b-%Y %H:%M", tz="UTC") mds$Year <- as.integer(format(mds$date, "%Y")) mds$Month <- format(mds$date, "%b") mds$DayTime <- format(mds$date, "%d, %R") tempds[[length(tempds) + 1]] <- mds months_years <- c(months_years, paste(month, '-', year)) } } }
错误根源
base_url写法错误:混用单引号和未解析的变量,导致本身就是无效字符串- URL拼接时直接使用整个
last_day向量,生成了长度为12的URL列表,而GET()函数仅接受单个URL
修复后的代码
# 初始化空列表和向量 tempds <- list() months_years <- character(0) years <- c('2022', '2023') months <- c('Jan','Feb', 'Mar', 'Apr','May','June','July','Aug','Sept','Oct','Nov','Dec') last_day <- c('31','28','31','30','31','30','31','31','30','31','30','31') # 循环遍历年份和月份 for (year in years){ for(month_idx in seq_along(months)) { month <- months[month_idx] day <- last_day[month_idx] print(paste(month, year)) # 拆分构造日期参数,避免拼接混乱 date_from <- paste0("01-", month, "-", year, "+00%3A00") date_to <- paste0(day, "-", month, "-", year, "+23%3A59") url <- paste0("https://www.smartgriddashboard.com/DashboardService.svc/data?area=generationactual®ion=ALL&datefrom=", date_from, "&dateto=", date_to) cat(url, "\n") # 终止条件:2023年9月停止 if (year == '2023' && month == 'Sept'){ print(paste(month, year, '已到达,终止循环')) break } # 发送GET请求并读取CSV数据 response <- GET(url) if (http_status(response)$status_code == 200) { content <- content(response, "text") mds <- read_csv(content, na = "-") # 处理数据 mds$date <- as.POSIXct(mds$`DATE & TIME`, format="%d-%b-%Y %H:%M", tz="UTC") mds$Year <- as.integer(format(mds$date, "%Y")) mds$Month <- format(mds$date, "%b") mds$DayTime <- format(mds$date, "%d, %R") tempds[[length(tempds) + 1]] <- mds months_years <- c(months_years, paste(month, '-', year)) } } }
修复说明
- 重构URL拼接逻辑:拆分
datefrom和dateto参数,避免字符串拼接混乱 - 通过索引匹配对应月份的最后一天,确保每次生成单个有效URL
- 修正终止条件位置,确保在发送请求前判断是否终止循环
内容的提问来源于stack exchange,提问作者Magnetar
相关产品推荐
相关产品推荐

