You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用str_locate时遇正则语法错误及越南语乱码问题求助

问题与解决方案

核心问题

  1. 越南语字符因编码问题显示为乱码(如Tên doanh nghi???p:),导致正则模式无效,触发Syntax error in regex pattern错误。
  2. 依赖str_locate+str_sub的文本位置提取方式脆弱,易受页面结构变化影响。

解决步骤

1. 修复字符编码

乱码源于网页读取或脚本文件编码不匹配:

  • 读取网页时强制指定UTF-8编码:
    # 修正搜索结果页读取
    search.result = read_html(link, encoding = "UTF-8")
    # 修正企业详情页读取
    com.page = read_html(paste0("https:", com.link), encoding = "UTF-8")
    
  • 确保R脚本文件编码为UTF-8(RStudio可通过右下角编码选项设置),避免硬编码的越南语标签出现乱码。

2. 替换为正确的越南语正则模式

将代码中乱码的标签替换为准确的越南语短语,示例:

乱码模式正确越南语标签
Tên doanh nghi???p:Tên doanh nghiệp:
Mã s??? thu???:Mã số thuế:
Tình tr???ng ho???t Tình trạng hoạt động:
???a ch???:Địa chỉ:
i???n tho???i:Điện thoại:

3. 改用更可靠的元素定位提取

放弃依赖文本位置的方式,直接通过HTML结构定位字段,稳定性更强:

方法1:XPath兄弟节点定位

# 提取企业名称
com.name = com.page %>% 
  html_nodes(xpath = '//*[contains(text(), "Tên doanh nghiệp:")]/following-sibling::div') %>% 
  html_text() %>% trim()

# 提取税号
com.mst = com.page %>% 
  html_nodes(xpath = '//*[contains(text(), "Mã số thuế:")]/following-sibling::div') %>% 
  html_text() %>% trim()

# 其他字段同理替换对应标签

方法2:批量提取所有信息项

# 获取所有信息项节点
info_items = com.page %>% html_nodes(".company-info-item")

# 批量解析标签和值
info_df = lapply(info_items, function(item) {
  label = item %>% html_node("span") %>% html_text()
  value = item %>% html_node("div") %>% html_text() %>% trim()
  data.frame(label = label, value = value, stringsAsFactors = FALSE)
}) %>% do.call(rbind, .)

# 按需提取字段
com.name = info_df$value[info_df$label == "Tên doanh nghiệp:"]
com.mst = info_df$value[info_df$label == "Mã số thuế:"]

4. 优化数据存储效率

避免循环中反复rbind,改用列表存储后统一合并:

# 初始化空列表
info_list = list()
count = 1

for (mst in list.mst) {
  link = paste0(link.source, mst,'/')
  message(link)
  search.result = read_html(link, encoding = "UTF-8") 
  all.com.link = search.result %>% html_nodes(".company-name a") %>% html_attr('href') %>% unique()
  
  for (com.link in all.com.link) {
    com.page = read_html(paste0("https:", com.link), encoding = "UTF-8")
    # ... 提取所有字段 ...
    
    # 将单条企业信息存入列表
    info_list[[count]] = data.frame(
      com.name = com.name,
      com.mst = com.mst,
      com.active = com.active,
      com.add = com.add,
      com.tel = com.tel,
      com.cap.phep = com.cap.phep,
      com.hoat.dong = com.hoat.dong,
      com.nganh = com.nganh,
      stringsAsFactors = FALSE
    )
    count = count + 1
  }
}

# 合并所有数据
info = do.call(rbind, info_list)

内容的提问来源于stack exchange,提问作者Tann

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 11:25:43