如何将NIBRS的ASCII格式2021年逮捕数据转为可用数据集?
处理NIBRS固定宽度ASCII数据:R与Stata导入指南
这类ASCII文件属于固定宽度格式,每个字段对应字符串的特定位置,只需按照NIBRS帮助文件里的字段位置定义,就能拆分出可用数据集。以下是R和Stata的具体操作方法:
R 中的导入方法
方法1:基础R的read.fwf()
适合小数据集,直接用列宽度定义拆分:
# 第一步:根据NIBRS帮助文件,定义每个字段的宽度(比如字段占1-10位,宽度就是10) col_widths <- c(10, 3, 4, 5, 7, 1, 9, 25, 10) # 示例宽度,替换为实际字段宽度 col_names <- c("AgencyID", "Year", "Month", "OffenseCode", "SomeCol", "Flag", "ArrestNum", "City", "State") # 对应字段名 # 第二步:读取文件,自动拆分 nibrs_data <- read.fwf( "nibrs_arrests_2021.txt", widths = col_widths, col.names = col_names, stringsAsFactors = FALSE, strip.white = TRUE # 自动去除字段前后的空白填充 )
方法2:tidyverse的read_fwf()
适合大数据集,速度更快,支持直接用起始/结束位置定义:
library(readr) # 第一步:根据帮助文件,定义每个字段的起始和结束位置 col_positions <- fwf_positions( start = c(1, 11, 14, 18, 23, 30, 31, 40, 65), # 示例起始位 end = c(10, 13, 17, 22, 29, 30, 39, 64, 74), # 示例结束位 col_names = c("AgencyID", "Year", "Month", "OffenseCode", "SomeCol", "Flag", "ArrestNum", "City", "State") ) # 第二步:读取文件 nibrs_data <- read_fwf( "nibrs_arrests_2021.txt", col_positions = col_positions, trim_ws = TRUE # 清除字段前后空白 )
Stata 中的导入方法
方法1:直接用infix命令
一行命令完成,适合字段不多的情况:
// 按位置定义变量:字符型加str,数字型直接写变量名和位置 infix str AgencyID 1-10 Year 11-13 Month 14-17 OffenseCode 18-22 SomeCol 23-29 Flag 30 ArrestNum 31-39 str City 40-64 str State 65-74 using "nibrs_arrests_2021.txt", clear
- 说明:
clear表示清除当前内存中的数据;如果要测试前10行,加in 1/10即可。
方法2:使用字典文件(复杂结构首选)
如果字段数量多,先写一个字典文件(比如命名为nibrs_dict.dct):
dictionary using "nibrs_arrests_2021.txt" { str AgencyID 1-10 int Year 11-13 int Month 14-17 int OffenseCode 18-22 int SomeCol 23-29 str Flag 30 long ArrestNum 31-39 str City 40-64 str State 65-74 }
然后在Stata中执行:
use "nibrs_dict.dct", clear
关键注意事项
- 必须严格对应NIBRS帮助文件里的字段位置,位置错一位就会导致所有数据错位。
- 测试时优先导入少量行验证拆分结果,没问题再处理全量数据。
- 字符型字段的空白填充可以用工具自动清除,避免后续分析出错。
内容的提问来源于stack exchange,提问作者leecarvallo
相关产品推荐
相关产品推荐

