在R中基于指定地址获取Block Group普查数据的问题求助
问题:R语言提取普查区块组(Block Group)数据失败
我在使用R语言提取Block Group(区块组)的普查数据时遇到问题。手里有一批具体地址,需要先通过地址匹配到对应的Block Group,再提取该区块组的白人/黑人/亚裔/西班牙裔/两种族/其他种族占比等普查数据,但代码无法正常运行,求帮忙解决。
报错信息
Error in get_acs(...) : unused argument (blockgroup = block_group)
原代码
library(httr) library(tidycensus) library(purrr) get_geocode_data <- function(address) { # The URL for the geocoding request url <- "https://geocoding.geo.census.gov/geocoder/geographies/onelineaddress" # The parameters for the request params <- list( address = address, format = "json", benchmark = "Public_AR_Current", vintage = "Current_Current" ) # Make the request response <- GET(url, query = params) # Parse the response content <- content(response, "parsed", "application/json") # Extract the geographies data if it exists if (length(content$result$addressMatches) > 0) { geographies <- content$result$addressMatches[[1]]$geographies } else { geographies <- NULL } return(geographies) } addresses <- c( "Administration Building, 400 Bizzell St, College Station, TX 77843", "1210 Varsity Dr., E. Carroll Joyner Visitor Center, Raleigh, NC 27606", "1331 Cir Park Dr, Knoxville, TN 37916" ) geocoding_data <- map(addresses, get_geocode_data) # Define census API key census_api_key("94d96c1dfb04afe285d0db958f78ded8a4b197f7") get_population <- function(geographies) { # If geographies is NULL, return NA if (is.null(geographies)) { return(NA) } # Extract the state, county, tract, and block group state <- geographies$`2010 Census Blocks`[[1]]$STATE county <- geographies$`2010 Census Blocks`[[1]]$COUNTY tract <- geographies$`2010 Census Blocks`[[1]]$TRACT block_group <- geographies$`2010 Census Blocks`[[1]]$BLKGRP # Print out the geography parameters print(paste("State:", state)) print(paste("County:", county)) print(paste("Tract:", tract)) print(paste("Block group:", block_group)) # Get the total population for the block group population <- get_acs( geography = "block group", variables = "B01003_001", state = state, county = county, tract = tract, blockgroup = block_group ) return(population$estimate) } # Get the total population for each block group populations <- map(geocoding_data, get_population)
问题原因
报错直接原因是get_acs函数没有blockgroup这个参数,正确参数名是block_group(带下划线)。另外原代码只获取了总人口,没覆盖你需要的种族占比数据,下面是修正后的完整代码。
修正后的代码
library(httr) library(tidycensus) library(purrr) library(dplyr) library(tidyr) # 地理编码函数:直接获取地址对应的Block Group信息 get_geocode_data <- function(address) { url <- "https://geocoding.geo.census.gov/geocoder/geographies/onelineaddress" params <- list( address = address, format = "json", benchmark = "Public_AR_Current", vintage = "Current_Current" ) response <- GET(url, query = params) content <- content(response, "parsed", "application/json") if (length(content$result$addressMatches) > 0) { # 直接提取区块组层级数据,不用从Block层级转 geographies <- content$result$addressMatches[[1]]$geographies$`Census Block Groups`[[1]] } else { geographies <- NULL } return(geographies) } # 待处理地址列表 addresses <- c( "Administration Building, 400 Bizzell St, College Station, TX 77843", "1210 Varsity Dr., E. Carroll Joyner Visitor Center, Raleigh, NC 27606", "1331 Cir Park Dr, Knoxville, TN 37916" ) # 获取每个地址的地理编码数据 geocoding_data <- map(addresses, get_geocode_data) # 设置Census API Key(建议执行一次install=TRUE,之后不用重复写) census_api_key("94d96c1dfb04afe285d0db958f78ded8a4b197f7", install = TRUE) # 提取区块组的种族人口及占比数据 get_race_data <- function(geographies) { if (is.null(geographies)) { return(tibble( address = NA, total_pop = NA, white_pct = NA, black_pct = NA, asian_pct = NA, hispanic_pct = NA, two_or_more_pct = NA, other_race_pct = NA )) } # 提取地理参数 state <- geographies$STATE county <- geographies$COUNTY tract <- geographies$TRACT block_group <- geographies$BLKGRP # 定义需要的种族/族裔变量(ACS 5年估计,2021年为例) race_vars <- c( total_pop = "B01003_001", # 总人口 white = "B02001_002", # 非西班牙裔白人 black = "B02001_003", # 非西班牙裔黑人 asian = "B02001_005", # 非西班牙裔亚裔 hispanic = "B03003_003", # 西班牙裔(族裔,非种族) two_or_more = "B02001_008", # 两种或以上种族 other_race = "B02001_004" # 其他单一种族 ) # 获取ACS数据,修正block_group参数名 acs_data <- get_acs( geography = "block group", variables = race_vars, state = state, county = county, tract = tract, block_group = block_group, year = 2021 ) # 整理数据为宽格式并计算占比 result <- acs_data %>% select(variable, estimate) %>% pivot_wider(names_from = variable, values_from = estimate) %>% mutate( white_pct = round(white / total_pop * 100, 2), black_pct = round(black / total_pop * 100, 2), asian_pct = round(asian / total_pop * 100, 2), hispanic_pct = round(hispanic / total_pop * 100, 2), two_or_more_pct = round(two_or_more / total_pop * 100, 2), other_race_pct = round(other_race / total_pop * 100, 2), address = addresses[which(map_lgl(geocoding_data, ~ identical(.x, geographies)))] ) %>% select(address, total_pop, ends_with("_pct")) return(result) } # 批量获取所有地址的种族数据 race_results <- map_dfr(geocoding_data, get_race_data) # 查看最终结果 print(race_results)
额外说明
- 地理编码优化:直接提取
Census Block Groups层级数据,避免从Block层级转取参数的冗余和错误。 - 参数修正:将
blockgroup改为block_group,匹配get_acs官方参数要求。 - 变量补充:覆盖了你需要的所有种族/族裔变量,注意西班牙裔是独立的族裔分类,不属于种族变量范畴。
- 数据整理:自动计算各群体占比并整理为易读的宽格式。
- API安全:建议执行
census_api_key("你的密钥", install = TRUE)一次,密钥会被保存到本地配置文件,后续脚本无需重复编写。
内容的提问来源于stack exchange,提问作者Fox_Summer
相关产品推荐
相关产品推荐

