You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在无命名空间XML中分别提取英文与日文字体名称?

问题描述

需要处理以下无命名空间的XML文件:

<?xml version="1.0" encoding="UTF-8"?>
<typekitSyncState>
  <state>2f1b61f7296e340f27be2925b5485719f04efcb8</state>
  <fonts type="array">
    <font>
      <url>https://someURI</url>
      <id>25367</id>
      <properties>
        <fullName>Heisei Kaku Gothic Std W3</fullName>
        <familyName>Heisei Kaku Gothic Std</familyName>
        <variationName>W3</variationName>
        <familyURL>https://typekit.com/fonts/heisei-kaku-gothic-std</familyURL>
        <familyWebId>mpll</familyWebId>
        <fvd>n3</fvd>
        <isVariable>false</isVariable>
        <i18n>
          <locales type="array">
            <locale>
              <ianaTag>ja</ianaTag>
              <fullName>平成角ゴシック Std W3</fullName>
              <familyName>平成角ゴシック Std</familyName>
            </locale>
          </locales>
        </i18n>
      </properties>
    </font>
  </fonts>
</typekitSyncState>

现有代码能提取所有fullName,但无法区分英文和日文本地化名称:

for element in root.iter():
  if element.tag == "fullName":
    print("%s - %s" % (element.tag, element.text))

输出:

fullName - Heisei Kaku Gothic Std W3
fullName - 平成角ゴシック Std W3

尝试多层循环解析时,if child4.tag == "fullName"判断不生效,希望实现如下逻辑:分别提取英文全称、家族名、变体名及日文全称,输出类似:

Heisei Kaku Gothic Std W3, Heisei Kaku Gothic Std, W3, 平成角ゴシック Std W3
解决方案

可以利用ElementTree的节点定位能力,直接按层级提取对应字段,避免无差别遍历:

import xml.etree.ElementTree as ET

# 解析XML文件(若为字符串则用ET.fromstring(xml_content))
tree = ET.parse('your_xml_file.xml')
root = tree.getroot()

# 遍历每个font节点
for font in root.findall('.//font'):
    # 提取properties下的英文名称字段
    properties = font.find('properties')
    en_fullname = properties.findtext('fullName', default='')
    en_familyname = properties.findtext('familyName', default='')
    variation_name = properties.findtext('variationName', default='')
    
    # 提取日文本地化全称
    ja_fullname = ''
    locale = properties.find('.//locale')
    if locale is not None:
        ja_fullname = locale.findtext('fullName', default='')
    
    # 输出结果
    print(f"{en_fullname}, {en_familyname}, {variation_name}, {ja_fullname}")

代码说明

  • 用findall('.//font')直接定位所有font节点,避免冗余的多层嵌套循环
  • 通过findtext方法精准获取节点文本,同时设置默认值防止节点缺失导致报错
  • 利用XPath表达式.//locale快速定位本地化节点,判断存在后提取日文全称

运行代码后输出:

Heisei Kaku Gothic Std W3, Heisei Kaku Gothic Std, W3, 平成角ゴシック Std W3

若使用BeautifulSoup,也可通过层级定位实现:

from bs4 import BeautifulSoup

xml_content = """[你的XML字符串内容]"""
soup = BeautifulSoup(xml_content, 'xml')

for font in soup.find_all('font'):
    properties = font.properties
    en_fullname = properties.fullName.text if properties.fullName else ''
    en_familyname = properties.familyName.text if properties.familyName else ''
    variation_name = properties.variationName.text if properties.variationName else ''
    
    ja_fullname = ''
    locale = properties.find('locale')
    if locale:
        ja_fullname = locale.fullName.text if locale.fullName else ''
    
    print(f"{en_fullname}, {en_familyname}, {variation_name}, {ja_fullname}")

内容的提问来源于stack exchange,提问作者Ghoul Fool

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 10:40:04