You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R进行HTML数据爬取:如何为爬取的国家分配所属大洲

How to Map Scraped Countries to Their Continents in R

Hey there! Let's tackle this problem of matching your scraped countries to their continents—here are a few solid approaches depending on your situation:

1. Scrape Continent Data Directly from the Source Website

Chances are the website you're pulling countries from already groups them by continent in its HTML structure. This is the most reliable method because it uses the exact categorization from your target site.

First, use your browser's developer tools (F12) to inspect how the page organizes continents and countries. For example, if each continent is wrapped in a section with a class like continent-group, and the continent name is in a heading tag (e.g., h2 or h3), you can extract both in one go:

# Load purrr if you haven't already (it's great for iterative scraping)
library(purrr)

# Grab all continent sections from the page (adjust selector to match your site)
continent_sections <- website1 %>% html_nodes(".continent-group")

# For each section, extract continent name and its countries, then combine into a dataframe
country_continent_df <- map_dfr(continent_sections, function(section) {
  # Get the continent name from the section's heading (update selector to your site's tag)
  continent <- section %>% html_node("h3") %>% html_text(trim = TRUE)
  # Get all countries in this section (reuse your existing selector)
  countries <- section %>% html_nodes(".miscTxt a") %>% html_text(trim = TRUE)
  # Return a tibble for this continent
  tibble(Country = countries, Continent = continent)
})

# View the result
country_continent_df

Just tweak the CSS selectors (.continent-group, h3) to match the actual structure of your target website—this will give you a perfect 1:1 match between countries and their continents as defined by the site.

2. Use a Pre-built Country-to-Continent Mapping Package

If scraping directly from the site isn't feasible (e.g., the page doesn't group countries by continent), the countrycode package is a fantastic tool for matching country names to standard continent classifications.

# Install and load the package if you haven't
install.packages("countrycode")
library(countrycode)

# Convert your country list to a dataframe with continent matches
country_continent_df <- tibble(Country = countries_list) %>%
  mutate(
    Continent = countrycode(
      sourcevar = Country,
      origin = "country.name",
      destination = "continent",
      warn = TRUE # This will alert you to any country names that don't match
    )
  )

Note: You might run into minor mismatches if the scraped country names use non-standard spellings (e.g., "USA" vs "United States"). The warn = TRUE flag will help you identify these, and you can clean up your country names manually or use the countrycode package's alias support to fix them.

3. Create a Manual Mapping Table

For edge cases where the above methods fail (e.g., very specific country names or regional groupings), you can build a custom mapping table:

# Create a tibble with your country-continent pairs
custom_continent_map <- tibble(
  Country = c("Country A", "Country B", "Country C"), # Replace with your scraped countries
  Continent = c("Asia", "Europe", "Africa") # Corresponding continents
)

# Join your scraped countries with the custom map
country_continent_df <- tibble(Country = countries_list) %>%
  left_join(custom_continent_map, by = "Country")

# Check for any unmatched countries (NA values in Continent column)
filter(country_continent_df, is.na(Continent))

This gives you full control over the categorization, which is useful if the target site uses non-standard continent groupings.

内容的提问来源于stack exchange,提问作者Hanish Kiran Sanghrajka

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 15:42:29