Python解析XSD文件时LineNumber字段值异常被覆盖的问题排查与解决咨询
Python解析XSD文件时LineNumber字段值异常被覆盖的问题排查与解决咨询
最近我在解析XSD文件时碰到了一个特别诡异的问题,实在摸不着头脑,来求助各位大佬:
先给大家看一下我要解析的XSD元素示例:
<xsd:element name="Example" type="ExampleType" minOccurs="0"> <xsd:annotation> <xsd:documentation> <Description>This is the description</Description> <LineNumber>4</LineNumber> </xsd:documentation> </xsd:annotation> </xsd:element>
我用Python的xml.etree.ElementTree库来做解析,核心代码如下:
# 依赖xml.etree.ElementTree库 annotation = element.find(".//xsd:annotation", namespace) if annotation is not None: documentation = annotation.find(".//xsd:documentation", namespace) if documentation is not None: for doc_child in documentation: tag = doc_child.tag.split('}')[-1] element_dict[element_name][tag] = doc_child.text.strip()
现在诡异的现象出现了:
- 直接打印解析后的
element_dict时,结果是完全正确的:
Description: This is the description LineNumber: 4
- 可一旦把
element_dict转成DataFrame再导出到Excel,LineNumber的值居然变成了:
LineNumber: Part A Line 12
我已经确认,这个“Part A Line 12”是XSD文件里后面另一个元素的LineNumber值,但我完全搞不懂:当前元素的解析逻辑为什么会拉取到后面元素的值?代码里到底哪里出了问题?另外我该怎么调整解析逻辑,才能正确获取当前元素对应的Description和LineNumber,避免被后面的元素值覆盖呢?
备注:内容来源于stack exchange,提问作者Broderick Boucher
相关产品推荐
相关产品推荐

