如何用Beautiful Soup提取XML中的非空标签有效信息?
提取Beautiful Soup Entry标签中的有效内容
给定一组Beautiful Soup的<Entry>标签元素列表,需遍历每个Entry,提取所有包含有效信息的标签(忽略无内容的空标签,比如带有xsi:nil="true"属性或文本/子标签为空的标签)。
需提取的目标内容
- 第一个Entry需提取:
<DateFormattedForTHForm>07/01/2022</DateFormattedForTHForm> <DateFormattedForTHForm>07/01/2023</DateFormattedForTHForm> <FormDescription>Notification Of Settlement</FormDescription> <FormNumber>WC 99 06 04</FormNumber> - 第二个Entry需提取:
<DisplayName>Mallesham Yamulla</DisplayName> <FEINOrSSN>123-45-6789</FEINOrSSN> <formsMaskedSSN_and_NoMaskFEIN>**-***-8834</formsMaskedSSN_and_NoMaskFEIN> <PrimaryAddress> <AddressLine1>A</AddressLine1> <AddressLine123>B</AddressLine123> <CityStateZip>ENID, OK 73703</CityStateZip> <Country>IND</Country> </PrimaryAddress>
实现代码
from bs4 import BeautifulSoup, Tag def extract_valid_tags(element): """递归提取元素下的有效标签,忽略空标签""" valid_tags = [] for child in element.children: if not isinstance(child, Tag): continue # 跳过非标签类型的节点 # 过滤带xsi:nil="true"的空标签 if child.get('xsi:nil') == 'true': continue # 判断标签是否有有效内容:文本非空 或 存在有效子标签 has_valid_text = child.get_text(strip=True) != '' has_valid_kids = len(extract_valid_tags(child)) > 0 if has_valid_text or has_valid_kids: # 清理标签内部,仅保留有效子标签 cleaned_tag = BeautifulSoup(child.prettify(), 'xml').find(child.name) cleaned_tag.contents = extract_valid_tags(child) valid_tags.append(cleaned_tag) return valid_tags # 假设entries是你的Entry标签列表 for idx, entry in enumerate(entries, 1): print(f"第{idx}个Entry提取结果:") for tag in extract_valid_tags(entry): print(tag.prettify())
代码逻辑说明
- 递归遍历每个标签的子元素,跳过文本、注释等非标签节点
- 直接排除带有
xsi:nil="true"属性的空标签 - 仅保留两类标签:自身文本内容非空的标签,或包含有效子标签的父标签
- 对符合条件的标签,清理其内部结构,只保留有效子标签,确保输出结果简洁
内容的提问来源于stack exchange,提问作者myamulla_ciencia
相关产品推荐
相关产品推荐

