求助:使用R xml2库提取XBRL文件节点失败问题排查
Hey there, let's work through why your xml2 functions aren't playing nice with that XBRL file! XBRL isn't just standard XML—it has unique structural quirks, especially around namespaces, that almost always cause this kind of issue. Let's break down the fixes step by step.
The #1 Culprit: XBRL Namespaces
XBRL relies heavily on XML namespaces (think of them as "prefixes" for element categories like US GAAP tags, core XBRL structures, etc.). xml2 doesn't auto-detect or apply these namespaces by default, so trying to query elements with just their raw tag name (like //Assets) will return nothing.
Fix: Explicitly Use Namespaces in Queries
First, let's map out all the namespaces in your file, then use those prefixes in your XPath queries:
library(xml2) library(tidyverse) # Read the file (you already have this part) eqr <- read_xml(x = "https://www.sec.gov/Archives/edgar/data/906107/000156459018010687/eqr-20180331.xml") # Extract and view all namespaces in the document ns <- xml_ns(eqr) print(ns)
You'll see output like us-gaap: http://xbrl.us/us-gaap/2018-01-31 or xbrli: http://www.xbrl.org/2003/instance—these are the prefixes you need to attach to element names.
Now use these prefixes to target elements properly:
# Example: Extract all US GAAP "Assets" elements total_assets <- xml_find_all(eqr, "//us-gaap:Assets", ns) %>% xml_text() %>% as.numeric() print(total_assets) # Example: Replace your xml_child(eqr, 3) with a reliable namespace-based query for units units <- xml_find_all(eqr, "//xbrli:unit", ns) View(units)
Other Common Pitfalls
Avoid relying on position-based child queries
xml_child(eqr, 3)works for your unit node now, but XBRL structure can vary between filings. Using XPath with namespaces is far more reliable than hardcoding positions.Check element hierarchy with
xml_structure()
If you're still stuck, visualize the document's structure to confirm where your target elements live:xml_structure(eqr)This will show you the nested hierarchy and which namespaces apply to each level.
Don't forget context/unit references
XBRL data often links values to context (date/entity) and unit (currency/measure) nodes via ID attributes. To get complete, meaningful data, you'll need to match these references. For example:# Grab a value and its associated context ID asset_value <- xml_find_first(eqr, "//us-gaap:Assets", ns) context_id <- xml_attr(asset_value, "contextRef") # Find the matching context node to get the reporting date context <- xml_find_first(eqr, paste0("//xbrli:context[@id='", context_id, "']"), ns) reporting_date <- xml_find_first(context, ".//xbrli:endDate", ns) %>% xml_text()
内容的提问来源于stack exchange,提问作者Chris Kiniry

