如何用BeautifulSoup提取XML中非嵌套的pos标签值?
提取BeautifulSoup解析XML时的非嵌套
<pos>标签值 原始XML结构
<?xml version="1.0" encoding="UTF-8" ?> <main_heading timestamp="20220113"> <details> <offer id="11" parent_id="12"> <name>Alpha</name> <pos>697</pos> <kat_pis> <pos kat="2">112</pos> </kat_pis> </offer> <offer id="12" parent_id="31"> <name>Beta</name> <pos>099</pos> <kat_pis> <pos kat="2">113</pos> </kat_pis> </offer> </details> </main_heading>
当前代码及问题
原代码会提取所有<pos>标签内容,包括嵌套在<kat_pis>里的:
soup = BeautifulSoup(file, 'xml') pos = [] for i in (soup.find_all('pos')): pos.append(i.text)
得到结果:['697', '112', '099', '113'],但只需要直接属于<offer>下的非嵌套<pos>值,即['697', '099']。
解决方案
方案1:直接定位<offer>下的直接子<pos>标签
先遍历所有<offer>标签,再在每个<offer>内只找直接子节点的<pos>:
soup = BeautifulSoup(file, 'xml') pos = [] for offer in soup.find_all('offer'): direct_pos = offer.find('pos', recursive=False) if direct_pos: pos.append(direct_pos.text)
方案2:过滤父标签非<kat_pis>的<pos>
遍历所有<pos>标签,判断其父标签是否不是<kat_pis>:
soup = BeautifulSoup(file, 'xml') pos = [] for pos_tag in soup.find_all('pos'): if pos_tag.parent.name != 'kat_pis': pos.append(pos_tag.text)
方案3:用CSS选择器直接筛选
利用CSS选择器的直接子元素语法>,精准选中<offer>的直接子<pos>:
soup = BeautifulSoup(file, 'xml') pos = [tag.text for tag in soup.select('offer > pos')]
以上三种方法都能得到期望的结果['697', '099']。
内容的提问来源于stack exchange,提问作者x89
相关产品推荐
相关产品推荐

