You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup从HTML中提取SVG元素与标题?

简洁提取SVG标题与路径坐标的方法

你可以直接在原函数中完成标题文本和路径坐标的提取,返回结构化数据,代码更紧凑实用:

修改后的函数(返回结构化列表)

from bs4 import BeautifulSoup

def scrape_spatial_data(page):
    html = page.inner_html("body")
    soup = BeautifulSoup(html, "html.parser")
    # 遍历每个g.myclass元素,提取标题和路径d属性
    return [
        {
            "title": g.title.get_text(strip=True),
            "path_coords": g.path["d"]
        }
        for g in soup.select("g.myclass")
        if g.title and g.path  # 可选:过滤缺失标题或路径的无效元素
    ]

说明

  • 利用BeautifulSoup的属性访问特性(g.title等价于g.find("title"))简化代码
  • get_text(strip=True)自动去除标题文本的首尾空白字符
  • g.path["d"]直接获取path标签的路径坐标属性值
  • 可选的条件判断可以避免因元素缺失导致的报错

如果处理大量数据,用生成器返回更节省内存:

def scrape_spatial_data(page):
    html = page.inner_html("body")
    soup = BeautifulSoup(html, "html.parser")
    for g in soup.select("g.myclass"):
        if g.title and g.path:
            yield {
                "title": g.title.get_text(strip=True),
                "path_coords": g.path["d"]
            }

内容的提问来源于stack exchange,提问作者felix1k

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 17:14:52