You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup提取嵌套XML中Soccer-Hidden标签内的信息?

问题

我用BeautifulSoup解析嵌套XML时,已经成功提取<f>标签的class和data-id属性,但没法获取<text class='Soccer-Hidden'>标签内的<span class='Soccer-Key'>、<span class='Soccer-Value'>以及<p>标签里的内容(比如Ronaldo),请问要在现有代码里加什么内容才能实现?

XML代码

<f transform="translate(7,7)" class="SoccerPlayer SoccerPlayer-11 Team-Away  Outcome-Complete" data-id="8">
    <rect x="-15" y="-15" width="30" height="30" transform="rotate(0)" class="SoccerShape"></rect>
    <text x="0" y="7" text-anchor="middle" transform="translate(0,0)rotate(0)">11</text>
    <text class="Soccer-Hidden">
        <div>
            <h3>
                <span class="Soccer-Key">
            Suc passes
          </span>
                <span class="Soccer-Value">
            82
          </span>
            </h3>
            <p>
          Ronaldo
        </p>
        </div>
    </text>
</f>

现有代码

from bs4 import BeautifulSoup as bs
soup=bs(xml, "xml")
for pr in soup.find_all("f"):
    try:
        player = pr['class']
        time = pr['data-id']
    except:
        pass
    print(player,time)
解决方案

你可以在循环里针对每个<f>标签,先定位到<text class="Soccer-Hidden">标签,再从中提取目标元素的内容。注意用.strip()去除文本里的多余空格,修改后的代码如下:

from bs4 import BeautifulSoup as bs
xml = """<f transform="translate(7,7)" class="SoccerPlayer SoccerPlayer-11 Team-Away  Outcome-Complete" data-id="8">
    <rect x="-15" y="-15" width="30" height="30" transform="rotate(0)" class="SoccerShape"></rect>
    <text x="0" y="7" text-anchor="middle" transform="translate(0,0)rotate(0)">11</text>
    <text class="Soccer-Hidden">
        <div>
            <h3>
                <span class="Soccer-Key">
            Suc passes
          </span>
                <span class="Soccer-Value">
            82
          </span>
            </h3>
            <p>
          Ronaldo
        </p>
        </div>
    </text>
</f>"""
soup=bs(xml, "xml")
for pr in soup.find_all("f"):
    try:
        player_classes = pr['class']
        data_id = pr['data-id']
        # 定位到Soccer-Hidden的text标签
        hidden_text = pr.find("text", class_="Soccer-Hidden")
        if hidden_text:
            # 提取Soccer-Key的内容
            key = hidden_text.find("span", class_="Soccer-Key").get_text(strip=True)
            # 提取Soccer-Value的内容
            value = hidden_text.find("span", class_="Soccer-Value").get_text(strip=True)
            # 提取p标签里的球员名
            player_name = hidden_text.find("p").get_text(strip=True)
            
            print(f"球员类: {player_classes}, ID: {data_id}")
            print(f"{key}: {value}, 球员名: {player_name}")
        else:
            print(f"球员类: {player_classes}, ID: {data_id}, 无隐藏信息")
    except Exception as e:
        print(f"解析出错: {e}")

关键说明

  • 用pr.find("text", class_="Soccer-Hidden")从当前<f>标签下定位到隐藏的文本容器
  • 用.find()进一步定位目标<span>和<p>标签
  • 用.get_text(strip=True)获取并清理文本内容,去除首尾空格和换行

内容的提问来源于stack exchange,提问作者JayRSP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 12:46:14