反洗钱核查爬取OFAC网站遇check_form()错误求助
问题解决:OFAC制裁网站爬取的check_form错误
错误原因
你直接将read_html返回的页面对象传给html_form_set,但这个函数要求参数是单个由html_form()生成的表单对象,而非原始页面。
修正后的代码
#install.packages("robotstxt") library(robotstxt) site<-"https://sanctionssearch.ofac.treas.gov/" paths_allowed(site) #install.packages("rvest") library(rvest) url <- "https://sanctionssearch.ofac.treas.gov/" search_first_name <- "John" search_last_name <- "Doe" # 获取页面 page <- read_html(url) # 提取页面中的搜索表单(取第一个表单) search_form <- html_form(page)[[1]] # 填充表单字段:OFAC表单需分别填写名和姓 filled_form <- html_form_set( search_form, "ctl00$MainContent$txtFirstName" = search_first_name, "ctl00$MainContent$txtLastName" = search_last_name ) # 提交表单获取响应 response <- html_form_submit(filled_form) # 提取并处理结果 results <- response %>% html_nodes("div.resultName") %>% html_text(trim = TRUE) # 检查目标姓名是否在结果中 target_name <- paste(search_first_name, search_last_name) name_found <- target_name %in% results cat("搜索姓名是否存在:", name_found, "\n")
关键修正点
- 先通过
html_form(page)获取页面表单集合,再选择目标表单(这里用[[1]]取第一个,可根据页面实际结构调整) - OFAC的搜索表单需要分别填写名(
txtFirstName)和姓(txtLastName),原代码只填了姓会导致搜索不准确 - 提交表单时操作的是填充后的表单对象,而非原始页面
内容的提问来源于stack exchange,提问作者zachi
相关产品推荐
相关产品推荐

