You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R提取XML多条目数据时遇重复行ID错误求排查

问题分析与解决方案

首先,你遇到的Error: Duplicate identifiers for rows (2, 3, 4)主要源于两个核心问题:递归函数的返回值处理逻辑错误,以及不符合预期的spread操作(你的目标是生成长格式的DataFrame,而非宽格式)。另外还有一处空值处理的语法错误需要修正。

具体问题拆解

  • 递归函数的返回逻辑错误:原函数中,当存在子节点时调用递归后,没有将递归返回的结果合并到当前的DataFrame中,且递归调用时传递的是未更新的df,导致数据要么重复要么未被正确收集。
  • 空值处理代码失效:is.na(f)<-which(f == '');f这行代码语法逻辑错误,无法完成将空字符串转为NA的需求。
  • 不必要的spread操作:你的预期输出是每行对应一个字段-值对的长格式,但原代码用spread强行转为宽格式,这不仅违背需求,还会因为行索引重复触发报错。

修正后的代码

library(xml2)
library(plyr)

setwd('c:/temp/xml/t')

findchildren <- function(nodes) {
  # 初始化空的临时DataFrame,指定字符串类型避免因子转换
  df <- data.frame(fieldname = character(), contents = character(), stringsAsFactors = FALSE)
  
  # 处理叶子节点(无后代的节点)
  leaf_nodes <- nodes[sapply(nodes, function(x) length(xml_children(x)) == 0)]
  if (length(leaf_nodes) > 0) {
    xmlvalue <- xml_text(leaf_nodes)
    xmlname <- xml_name(leaf_nodes)
    # 获取父节点路径,反转后拼接成下划线分隔的字符串
    xmlpath <- sapply(leaf_nodes, function(x) {
      gsub(', ', '_', toString(rev(xml_name(xml_parents(x)))))
    })
    # 针对Reason下的Description,生成带值的唯一fieldname
    fieldname <- ifelse(xmlpath == 'MyDecision_Decision_DecisionReasons_Reason',
                       paste(xmlpath, xmlname, xmlvalue, sep = '_'),
                       paste(xmlpath, xmlname, sep = '_'))
    # 正确处理空值,将空字符串转为NA
    contents <- ifelse(xmlvalue == '', NA, xmlvalue)
    # 合并当前叶子节点的数据到临时df
    df <- rbind(df, data.frame(fieldname, contents, stringsAsFactors = FALSE))
  }
  
  # 处理非叶子节点,递归调用并合并子节点的结果
  non_leaf_nodes <- nodes[sapply(nodes, function(x) length(xml_children(x)) > 0)]
  if (length(non_leaf_nodes) > 0) {
    for (node in non_leaf_nodes) {
      child_df <- findchildren(xml_children(node))
      df <- rbind(df, child_df)
    }
  }
  
  return(df)
}

allxmldata <- data.frame(fieldname = character(), contents = character(), stringsAsFactors = FALSE)
inputfiles <- as.character('test.xml')

# 遍历处理所有XML文件
for (x in inputfiles) {
  xml_doc <- read_xml(x)
  df1 <- findchildren(xml_children(xml_doc))
  allxmldata <- rbind.fill(allxmldata, df1)
}

# 查看最终结果
print(allxmldata)

关键修改说明

  1. 重构递归逻辑:
    函数不再依赖外部传入的df参数,而是在内部初始化并通过递归调用合并子节点的结果,确保所有层级的节点数据都被正确收集,避免了参数传递导致的数据丢失或重复。
  2. 修正空值处理:
    使用ifelse(xmlvalue == '', NA, xmlvalue)替代原错误语法,准确完成空字符串转NA的需求。
  3. 移除spread操作:
    直接保留长格式的DataFrame结构,完全匹配你预期的输出样式,同时避免了宽格式转换带来的重复行报错。

最终预期输出

运行修正后的代码后,你会得到如下结构的DataFrame:

fieldnamecontents
MyDecision_Decision_DecisionIDX1234
MyDecision_Decision_DecisionReasons_Reason_Description_DOBMismatchDOBMismatch
MyDecision_Decision_DecisionReasons_Reason_Description_PrimaryChecksFailPrimaryChecksFail
MyDecision_Decision_DecisionReasons_Reason_Description_IncomeReferralIncomeReferral

内容的提问来源于stack exchange,提问作者Northernlad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:09:12