如何用BeautifulSoup从HTML中提取SVG元素与标题?
简洁提取SVG标题与路径坐标的方法
你可以直接在原函数中完成标题文本和路径坐标的提取,返回结构化数据,代码更紧凑实用:
修改后的函数(返回结构化列表)
from bs4 import BeautifulSoup def scrape_spatial_data(page): html = page.inner_html("body") soup = BeautifulSoup(html, "html.parser") # 遍历每个g.myclass元素,提取标题和路径d属性 return [ { "title": g.title.get_text(strip=True), "path_coords": g.path["d"] } for g in soup.select("g.myclass") if g.title and g.path # 可选:过滤缺失标题或路径的无效元素 ]
说明
- 利用BeautifulSoup的属性访问特性(
g.title等价于g.find("title"))简化代码 get_text(strip=True)自动去除标题文本的首尾空白字符g.path["d"]直接获取path标签的路径坐标属性值- 可选的条件判断可以避免因元素缺失导致的报错
如果处理大量数据,用生成器返回更节省内存:
def scrape_spatial_data(page): html = page.inner_html("body") soup = BeautifulSoup(html, "html.parser") for g in soup.select("g.myclass"): if g.title and g.path: yield { "title": g.title.get_text(strip=True), "path_coords": g.path["d"] }
内容的提问来源于stack exchange,提问作者felix1k
相关产品推荐
相关产品推荐

