使用Rvest提取列表中各城市对应URL的技术问题
问题
我正在研究rvest包,遇到了从列表中提取URL的问题。我的目标是生成一个包含Country(国家)、**City(城市)**和城市对应URL的dataframe(数据框),目前已拥有包含各国信息的dataframe以及对应每个国家的城市列表。
我的疑问是:如何定位每个城市以获取其对应的URL链接?我尝试获取wikitable sortable jquery-tablesorter类下td标签内的href属性,但运行代码links = webpage %>% html_node("href") %>% html_text()时仅得到主URL。
以下是我的代码:
# Get URL url = "https://en.wikipedia.org/wiki/List_of_towns_and_cities_with_100,000_or_more_inhabitants/country:_A-B" # Read the HTML code from the website page = read_html(url) # Get name of the countries countries = page %>% html_nodes(".mw-headline") %>% html_text() #Remove the last two items which are not countries countries = as.tibble(countries) %>% slice(1:(n()-2)) #Add row number to each Country to left_join later countries = rowid_to_column(countries, "column_label") # Get cities for that country # Still working on this since it includes the first table and I get blanks when I filter the html_nodes(".jquery-tablesorter td") tables = html_nodes(page, "table") tables = lapply(tables, html_table) #Remove fist element which is not a city, only on the first page tables = tables[-1] #---WIP # Get links for the cities, currently picks the main domain instead of the city # Can I add a clause before the html node to indicate I want the href from "wikitable sortable jquery-tablesorter"? links = page %>% html_attr("href") %>% html_text() #--- #Remove the Providence and Population columns and keeps City and URL tables = lapply(tables, "[", -c(2, 3)) #Standardize City as the column tables = map(tables, set_names, "City") # Flatten List all <- bind_rows(tables, .id = "column_label") %>% mutate(column_label = as.integer(column_label)) %>% left_join(countries, by = "column_label")
解决方案
你当前提取URL的代码逻辑有误,html_attr("href")直接作用在page上只会抓取页面根节点的href(也就是主域名相关),正确的做法是定位到每个国家对应的表格,再提取表格内城市链接的href属性。
分步修改:
- 定位目标表格并提取城市链接
直接筛选带有wikitable sortable jquery-tablesorter类的表格,然后遍历每个表格提取第一列(城市列)中的<a>标签的href属性:
# 获取所有目标表格(带指定类的表格,跳过第一个非城市表格) city_tables <- page %>% html_nodes("table.wikitable.sortable.jquery-tablesorter") # 遍历每个表格,提取城市名称和对应的URL city_data <- map_df(city_tables, function(table) { # 提取城市名称 cities <- table %>% html_nodes("td:first-child a") %>% html_text() # 提取城市对应的URL city_urls <- table %>% html_nodes("td:first-child a") %>% html_attr("href") # 组合成数据框 tibble(City = cities, URL = city_urls) }, .id = "column_label")
- 整合国家信息
你的国家提取逻辑可以保留,直接和上面的city_data合并即可:
# 合并国家和城市数据 final_df <- city_data %>% mutate(column_label = as.integer(column_label)) %>% left_join(countries, by = "column_label") %>% rename(Country = value) %>% # 重命名国家列 select(Country, City, URL) # 调整列顺序
完整修正后的代码:
library(rvest) library(dplyr) library(purrr) # Get URL url = "https://en.wikipedia.org/wiki/List_of_towns_and_cities_with_100,000_or_more_inhabitants/country:_A-B" # Read the HTML code from the website page = read_html(url) # Get name of the countries countries = page %>% html_nodes(".mw-headline") %>% html_text() # Remove the last two items which are not countries countries = as.tibble(countries) %>% slice(1:(n()-2)) # Add row number to each Country to left_join later countries = rowid_to_column(countries, "column_label") # 获取所有目标表格并提取城市名称和URL city_tables <- page %>% html_nodes("table.wikitable.sortable.jquery-tablesorter") city_data <- map_df(city_tables, function(table) { cities <- table %>% html_nodes("td:first-child a") %>% html_text() city_urls <- table %>% html_nodes("td:first-child a") %>% html_attr("href") tibble(City = cities, URL = city_urls) }, .id = "column_label") # 合并国家、城市、URL数据 final_df <- city_data %>% mutate(column_label = as.integer(column_label)) %>% left_join(countries, by = "column_label") %>% rename(Country = value) %>% select(Country, City, URL) # 查看结果 head(final_df)
关键说明:
- 使用
table.wikitable.sortable.jquery-tablesorter精准定位每个国家的城市表格,避免抓取无关表格; - 通过
td:first-child a定位到每个城市的链接标签,html_attr("href")提取链接路径,html_text()提取城市名称; - 用
map_df直接将每个表格的结果合并成一个数据框,简化后续的绑定操作。
内容的提问来源于stack exchange,提问作者Raul
相关产品推荐
相关产品推荐

