如何在R中每日向数据列表追加二手车均价数据
解决方案
步骤1:完善当日数据处理(加日期+清洗价格)
首先要把爬取到的文本格式价格转成数值,再给当日均价表加上日期列,用lubridate::today()自动获取当日日期:
library(rvest) library(tidyverse) library(lubridate) # 爬取页面数据 page <- read_html("*URL*") brand <- page %>% html_nodes("*css_for_car_brand*") %>% html_text() price_text <- page %>% html_nodes("*css_for_price*") %>% html_text() # 清洗价格:根据网站实际格式调整,比如移除货币符号、逗号 price <- price_text %>% str_remove_all("[^0-9.]") %>% as.numeric() # 生成带日期的当日均价宽表 daily_avg <- tibble(brand, price) %>% drop_na(price) %>% # 剔除价格缺失的无效数据 group_by(brand) %>% summarise(mean_price = mean(price, na.rm = TRUE)) %>% pivot_wider(values_from = mean_price, names_from = brand) %>% mutate(date = today()) %>% relocate(date) # 把日期列放在最前面
步骤2:合并并保存历史数据
通过判断历史数据文件是否存在,实现自动追加当日数据:
# 定义历史数据存储路径(用csv方便查看,也可以用RDS格式) history_file <- "used_car_avg_history.csv" if (file.exists(history_file)) { # 读取已有历史数据,合并当日数据 history_data <- read_csv(history_file, show_col_types = FALSE) updated_history <- bind_rows(history_data, daily_avg) } else { # 首次运行,直接用当日数据初始化历史表 updated_history <- daily_avg } # 保存更新后的历史数据 write_csv(updated_history, history_file)
可选:每日自动运行
如果需要定时执行脚本,Windows可以用taskscheduleR包,Linux/macOS用cronR包设置每日定时任务,比如每天早8点自动爬取更新。
注意点
- 价格清洗的正则要对应目标网站的格式,比如价格是"$12,345"或"¥5.6万",要调整
str_remove_all的规则 - 用
drop_na(price)避免因缺失值导致均值计算错误
内容的提问来源于stack exchange,提问作者Chrisabe
相关产品推荐
相关产品推荐

