如何在无命名空间XML中分别提取英文与日文字体名称?
问题描述
需要处理以下无命名空间的XML文件:
<?xml version="1.0" encoding="UTF-8"?> <typekitSyncState> <state>2f1b61f7296e340f27be2925b5485719f04efcb8</state> <fonts type="array"> <font> <url>https://someURI</url> <id>25367</id> <properties> <fullName>Heisei Kaku Gothic Std W3</fullName> <familyName>Heisei Kaku Gothic Std</familyName> <variationName>W3</variationName> <familyURL>https://typekit.com/fonts/heisei-kaku-gothic-std</familyURL> <familyWebId>mpll</familyWebId> <fvd>n3</fvd> <isVariable>false</isVariable> <i18n> <locales type="array"> <locale> <ianaTag>ja</ianaTag> <fullName>平成角ゴシック Std W3</fullName> <familyName>平成角ゴシック Std</familyName> </locale> </locales> </i18n> </properties> </font> </fonts> </typekitSyncState>
现有代码能提取所有fullName,但无法区分英文和日文本地化名称:
for element in root.iter(): if element.tag == "fullName": print("%s - %s" % (element.tag, element.text))
输出:
fullName - Heisei Kaku Gothic Std W3 fullName - 平成角ゴシック Std W3
尝试多层循环解析时,if child4.tag == "fullName"判断不生效,希望实现如下逻辑:分别提取英文全称、家族名、变体名及日文全称,输出类似:
Heisei Kaku Gothic Std W3, Heisei Kaku Gothic Std, W3, 平成角ゴシック Std W3
解决方案
可以利用ElementTree的节点定位能力,直接按层级提取对应字段,避免无差别遍历:
import xml.etree.ElementTree as ET # 解析XML文件(若为字符串则用ET.fromstring(xml_content)) tree = ET.parse('your_xml_file.xml') root = tree.getroot() # 遍历每个font节点 for font in root.findall('.//font'): # 提取properties下的英文名称字段 properties = font.find('properties') en_fullname = properties.findtext('fullName', default='') en_familyname = properties.findtext('familyName', default='') variation_name = properties.findtext('variationName', default='') # 提取日文本地化全称 ja_fullname = '' locale = properties.find('.//locale') if locale is not None: ja_fullname = locale.findtext('fullName', default='') # 输出结果 print(f"{en_fullname}, {en_familyname}, {variation_name}, {ja_fullname}")
代码说明
- 用
findall('.//font')直接定位所有font节点,避免冗余的多层嵌套循环 - 通过
findtext方法精准获取节点文本,同时设置默认值防止节点缺失导致报错 - 利用XPath表达式
.//locale快速定位本地化节点,判断存在后提取日文全称
运行代码后输出:
Heisei Kaku Gothic Std W3, Heisei Kaku Gothic Std, W3, 平成角ゴシック Std W3
若使用BeautifulSoup,也可通过层级定位实现:
from bs4 import BeautifulSoup xml_content = """[你的XML字符串内容]""" soup = BeautifulSoup(xml_content, 'xml') for font in soup.find_all('font'): properties = font.properties en_fullname = properties.fullName.text if properties.fullName else '' en_familyname = properties.familyName.text if properties.familyName else '' variation_name = properties.variationName.text if properties.variationName else '' ja_fullname = '' locale = properties.find('locale') if locale: ja_fullname = locale.fullName.text if locale.fullName else '' print(f"{en_fullname}, {en_familyname}, {variation_name}, {ja_fullname}")
内容的提问来源于stack exchange,提问作者Ghoul Fool
相关产品推荐
相关产品推荐

