导入含重复标签的XML:为每个P行保留父级HEAD信息
问题:XML解析生成包含父级HEAD的表格/DataFrame,每个
标签对应一行
需要解析一份XML文件,生成表格或DataFrame,要求:
- 每个
<P>标签对应表格的一行 - 每行需包含该
<P>的父级<DIV6>的<HEAD>内容、父级<DIV8>的<HEAD>内容
示例XML片段如下:
<DIV5 N="27" TYPE="PART" VOLUME="1" hierarchy_metadata="{"path":"/on/_SUBSTITUTE_DATE_/title-14/part-27","citation":"14 CFR Part 27","alternate_reference":"FAR Part 27"}"> <DIV6 N="A" TYPE="SUBPART" hierarchy_metadata="{"path":"/on/_SUBSTITUTE_DATE_/title-14/part-27/subpart-A","citation":"14 CFR Part 27 Subpart A","alternate_reference":"FAR Part 27 Subpart A"}"> <HEAD>Subpart A—General</HEAD> <DIV8 N="27.1" TYPE="SECTION" VOLUME="1" hierarchy_metadata="{"path":"/on/_SUBSTITUTE_DATE_/title-14/section-27.1","citation":"14 CFR 27.1","alternate_reference":"FAR 27.1"}"> <HEAD>§ 27.1 Applicability.</HEAD> <P>(a) This part prescribes airworthiness standards for the issue of type certificates, and changes to those certificates, for normal category rotorcraft with maximum weights of 7,000 pounds or less and nine or less passenger seats. </P> <P>(b) Each person who applies under Part 21 for such a certificate or change must show compliance with the applicable requirements of this part. </P> <P>(c) Multiengine rotorcraft may be type certified as Category A provided the requirements referenced in appendix C of this part are met. </P> <CITA TYPE="N">[Doc. No. 5074, 29 FR 15695, Nov. 24, 1964, as amended by Amdt. 27–33, 61 FR 21906, May 10, 1996; Amdt. 27–37, 64 FR 45094, Aug. 18, 1999] </CITA> </DIV8> <DIV8 N="27.2" TYPE="SECTION" VOLUME="1" hierarchy_metadata="{"path":"/on/_SUBSTITUTE_DATE_/title-14/section-27.2","citation":"14 CFR 27.2","alternate_reference":"FAR 27.2"}"> <HEAD>§ 27.2 Special retroactive requirements.</HEAD> <P>(a) For each rotorcraft manufactured after September 16, 1992, each applicant must show that each occupant's seat is equipped with a safety belt and shoulder harness that meets the requirements of paragraphs (a), (b), and (c) of this section. </P> <P>(1) Each occupant's seat must have a combined safety belt and shoulder harness with a single-point release. Each pilot's combined safety belt and shoulder harness must allow each pilot, when seated with safety belt and shoulder harness fastened, to perform all functions necessary for flight operations. There must be a means to secure belts and harnesses, when not in use, to prevent interference with the operation of the rotorcraft and with rapid egress in an emergency. </P> <P>(2) Each occupant must be protected from serious head injury by a safety belt plus a shoulder harness that will prevent the head from contacting any injurious object. </P> <P>(3) The safety belt and shoulder harness must meet the static and dynamic strength requirements, if applicable, specified by the rotorcraft type certification basis. </P> <P>(4) For purposes of this section, the date of manufacture is either— </P> <P>(i) The date the inspection acceptance records, or equivalent, reflect that the rotorcraft is complete and meets the FAA-Approved Type Design Data; or </P> <P>(ii) The date the foreign civil airworthiness authority certifies that the rotorcraft is complete and issues an original standard airworthiness certificate, or equivalent, in that country. </P> <P>(b) For rotorcraft with a certification basis established prior to October 18, 1999— </P> <CITA TYPE="N">[Doc. No. 26078, 56 FR 41051, Aug. 16, 1991, as amended by Amdt. 27–37, 64 FR 45094, Aug. 18, 1999] </CITA> </DIV8> </DIV6> </DIV5>
当前遇到的问题
使用Pandas的read_xml方法尝试解析时,只能获取到<DIV8>的<HEAD>内容和最后一个<P>标签的内容,其余<P>标签的数据全部丢失。核心问题是同一<DIV8>下存在多个同名的<P>标签,无法被正确识别并拆分。同时需要将<DIV6>的<HEAD>内容也作为列加入结果,不局限于使用Pandas。
用户尝试的代码:
import pandas as pd x=pd.read_xml("https://www.ecfr.gov/api/versioner/v1/full/2023-07-11/title-14.xml?part=27", xpath="/DIV5/DIV6//DIV8") x.to_csv("Export.csv")
注:已提前离线处理<P>标签中的斜体标记,无需考虑该问题。
解决方案
方案1:使用lxml手动解析(灵活可控)
通过遍历XML节点,手动提取所需数据并构建列表,最后转为DataFrame:
from lxml import etree import pandas as pd # 加载XML文件(可替换为本地文件路径) tree = etree.parse("https://www.ecfr.gov/api/versioner/v1/full/2023-07-11/title-14.xml?part=27") root = tree.getroot() data = [] # 遍历所有DIV6节点 for div6 in root.xpath("/DIV5/DIV6"): div6_head = div6.find("HEAD").text.strip() if div6.find("HEAD") is not None else "" # 遍历当前DIV6下的所有DIV8节点 for div8 in div6.xpath(".//DIV8"): div8_head = div8.find("HEAD").text.strip() if div8.find("HEAD") is not None else "" # 遍历当前DIV8下的所有P节点 for p in div8.xpath(".//P"): p_text = p.text.strip() if p.text is not None else "" data.append({ "DIV6_HEAD": div6_head, "DIV8_HEAD": div8_head, "P_CONTENT": p_text }) # 转为DataFrame并保存 df = pd.DataFrame(data) df.to_csv("Export.csv", index=False)
方案2:调整Pandas的read_xml参数(简化版)
通过修改xpath指向<P>节点,同时利用父节点引用提取上级HEAD:
import pandas as pd df = pd.read_xml( "https://www.ecfr.gov/api/versioner/v1/full/2023-07-11/title-14.xml?part=27", xpath="//P", converters={ "DIV6_HEAD": lambda x: x.strip(), "DIV8_HEAD": lambda x: x.strip(), "P_CONTENT": lambda x: x.strip() } ) # 提取父级DIV8和DIV6的HEAD df["DIV8_HEAD"] = df.apply(lambda row: row["parent"].find("HEAD").text.strip(), axis=1) df["DIV6_HEAD"] = df.apply(lambda row: row["parent"].getparent().find("HEAD").text.strip(), axis=1) # 保留所需列并重新排序 df = df[["DIV6_HEAD", "DIV8_HEAD", "text"]].rename(columns={"text": "P_CONTENT"}) df.to_csv("Export.csv", index=False)
说明:方案1兼容性更强,能处理更复杂的XML结构;方案2利用Pandas原生方法,代码更简洁,但依赖XML节点的层级关系稳定。
内容的提问来源于stack exchange,提问作者ScriptPhoenix
相关产品推荐
相关产品推荐

