使用rvest爬取网页时,变量值'F'被转为'FALSE'如何避免?
解决rvest爬取表格时'F'被转为FALSE的问题
rvest的html_table()函数默认开启自动类型转换(convert = TRUE),单字符的'F'会被误识别为逻辑值FALSE。可以通过以下两种方法避免这个问题:
方法一:关闭自动类型转换
直接在html_table()中设置convert = FALSE,所有单元格内容会以原始字符串形式保留:
call_link <- "https://www.cdisc.org/standards/therapeutic-areas/vaccines/vaccines-therapeutic-area-user-guide-v11/html" call_page <- read_html(call_link) call_table <- call_page %>% html_nodes("table.confluenceTable.dataset") %>% html_table(convert = FALSE) %>% .[[5]]
这种方法会让所有列变为字符类型,后续若需转换其他列的类型,可手动用type.convert()处理,比如:
call_table <- type.convert(call_table, as.is = c("需要保留字符的列名"))
方法二:手动提取单元格构建数据框
如果需要更精细的控制,可直接提取每个单元格文本再组装成数据框:
call_link <- "https://www.cdisc.org/standards/therapeutic-areas/vaccines/vaccines-therapeutic-area-user-guide-v11/html" call_page <- read_html(call_link) # 定位目标表格 target_table <- call_page %>% html_nodes("table.confluenceTable.dataset") %>% .[[5]] # 提取表头 headers <- target_table %>% html_nodes("th") %>% html_text() # 提取所有行的单元格内容 rows <- target_table %>% html_nodes("tr") %>% tail(-1) %>% # 跳过表头行 lapply(function(row) { row %>% html_nodes("td") %>% html_text() }) # 转换为数据框 call_table <- do.call(rbind.data.frame, rows) colnames(call_table) <- headers
这种方式完全由你控制每个单元格的内容,不会出现自动类型转换的问题。
内容的提问来源于stack exchange,提问作者jkatam
相关产品推荐
相关产品推荐

