如何在R中提取网页表格Your Choices列首个换行符前的内容
你可以借助tidyverse内置的stringr包的字符串处理函数实现需求,推荐使用str_replace方法,兼容性更强,即便某行内容没有换行符也不会返回空值。
完整调整后的代码
library(tidyverse) library(rvest) # read_html属于rvest包功能,需额外加载 content <- read_html("https://www.booking.com/hotel/mu/tamarin.en-gb.html?aid=356980&label=gog235jc-1DCAsonQFCE2hlcml0YWdlLWF3YWxpLWdvbGZIM1gDaJ0BiAEBmAExuAEXyAEM2AED6AEB-AECiAIBqAIDuAKiwqmEBsACAdICJGFkMTQ3OGU4LTUwZDMtNGQ5ZS1hYzAxLTc0OTIyYTRiZDIxM9gCBOACAQ&sid=729aafddc363c28a2c2c7379d7685d87&all_sr_blocks=36363601_246990918_2_85_0&checkin=2021-11-15&checkout=2021-11-20&dest_id=-1354779&dest_type=city&dist=0&from_beach_key_ufi_sr=1&group_adults=2&group_children=0&hapos=1&highlighted_blocks=36363601_246990918_2_85_0&hp_group_set=0&hpos=1&no_rooms=1&sb_price_type=total&sr_order=popularity&sr_pri_blocks=36363601_246990918_2_85_0__29200&srepoch=1619681695&srpvid=51c8354f03be0097&type=total&ucfs=1&req_children=0&req_adults=2&hp_refreshed_with_new_dates=1") tables <- content %>% html_table(fill = TRUE) second_table <- tables[[2]] # 处理Your Choices列,仅保留首个换行符前的内容 second_table <- second_table %>% mutate(`Your Choices` = str_replace(`Your Choices`, "\\n.*", "") %>% str_trim()) # 查看处理后结果 View(second_table)
代码说明
str_replace(Your Choices, "\\n.*", ""):匹配第一个换行符\n及之后的所有内容,替换为空字符串,仅保留换行前的文本str_trim():用来清除文本前后多余的空白字符,避免结果里出现不必要的空格、换行符
内容的提问来源于stack exchange,提问作者user3115933
相关产品推荐
相关产品推荐

