如何提取XML中无xml:lang属性的<sample>元素内容?
筛选XML中无xml:lang属性的元素
你的XML内容
<components version="1.0.0"> <component type="foo"> <sample>Foo</sample> <sample xml:lang="a">abc</sample> <sample xml:lang="b">efj</sample> </component> </components>
现有Python解析代码
from lxml import etree def parse(path: str): return etree.parse(path) def components(path: str) -> list: components = parse(path).xpath("/components/component") return list(components) def sample(path: str) -> str: sample = components(path)[0].find("sample").text return str(sample) path = "test.xml" print(sample(path))
实现方法
方法一:用XPath直接筛选(推荐)
利用XPath的条件表达式,直接定位没有xml:lang属性的<sample>元素,代码简洁高效:
from lxml import etree def parse(path: str): return etree.parse(path) def get_sample_without_lang(path: str) -> str: # XPath语法:not(@xml:lang) 表示元素不存在xml:lang属性 target_sample = parse(path).xpath("/components/component/sample[not(@xml:lang)]")[0] return target_sample.text path = "test.xml" print(get_sample_without_lang(path)) # 输出:Foo
方法二:遍历元素检查属性
如果需要遍历所有<sample>元素再判断,可以通过元素的attrib字典或get()方法检查属性是否存在:
from lxml import etree def parse(path: str): return etree.parse(path) def components(path: str) -> list: components = parse(path).xpath("/components/component") return list(components) def get_sample_without_lang(path: str) -> str: target_component = components(path)[0] # 遍历所有<sample>子元素 for sample_elem in target_component.findall("sample"): # 方式1:检查attrib字典中是否不含xml:lang if "xml:lang" not in sample_elem.attrib: return sample_elem.text # 方式2:用get方法,属性不存在时返回None # if sample_elem.get("xml:lang") is None: # return sample_elem.text return "" # 未找到符合条件的元素时返回空字符串 path = "test.xml" print(get_sample_without_lang(path)) # 输出:Foo
内容的提问来源于stack exchange,提问作者user22990699
相关产品推荐
相关产品推荐

