You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Ansible解析网站sitemap.xml并提取loc与priority字段

解决Ansible解析Sitemap.xml生成loc和priority字典列表的问题

问题背景

现有结构如下的sitemap.xml文件:

<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9" xmlns:xhtml="http://www.w3.org/1999/xhtml">
  <url>
    <loc>https://example.com/es/</loc>
    <lastmod>2023-02-15</lastmod>
    <changefreq>monthly</changefreq>
    <priority>0.5</priority>
  </url>
  <url>
    <loc>https://example/en/</loc>
    <lastmod>2023-02-15</lastmod>
    <changefreq>monthly</changefreq>
    <priority>0.5</priority>
  </url>
  <url>
    <loc>https://example.com/en/destinations/</loc>
    <lastmod>2021-09-16</lastmod>
    <changefreq>monthly</changefreq>
    <priority>0.5</priority>
  </url>
[...]
</urlset>

需通过Ansible完成以下操作:

  • 用ansible.builtin.uri下载文件并注册变量;
  • 遍历urlset内的url节点,生成包含loc和priority的字典列表;
  • 或转为JSON后完成同样操作。

当前使用community.general.xml模块未实现正确解析,parsedxml.xmlstring仍为原XML内容;ansible.netcommon.parse_xml过滤器因参数问题无法使用,需解决遍历节点生成目标字典列表的问题。


解决方案一:使用community.general.xml模块正确解析

调整XPath路径定位所有url节点,设置结构化输出参数,后续循环提取目标字段:

- name: Get the website 'sitemap.xml' file
  ansible.builtin.uri:
    url: "https://example.com/sitemap.xml"
    method: GET
    return_content: true
    headers:
      Accept: "application/xml"
    status_code: 200
    timeout: 5
  register: sitemap
  delegate_to: localhost

- name: Parse all url nodes from sitemap
  community.general.xml:
    xmlstring: "{{ sitemap.content }}"
    xpath: /s:urlset/s:url
    content: dictionary
    namespaces:
      s: http://www.sitemaps.org/schemas/sitemap/0.9
  register: parsed_urls
  delegate_to: localhost

- name: Generate list of loc-priority dictionaries
  set_fact:
    sitemap_entries: "{{ sitemap_entries | default([]) + [ {'loc': item.loc[0], 'priority': item.priority[0]} ] }}"
  loop: "{{ parsed_urls.matches }}"
  delegate_to: localhost

- name: Debug the result
  ansible.builtin.debug:
    var: sitemap_entries
  delegate_to: localhost

关键说明

  • xpath: /s:urlset/s:url 直接定位urlset下的所有url子节点;
  • content: dictionary 将每个url节点转为字典结构,子元素(如loc、priority)以列表形式存储值;
  • 通过loop遍历解析后的节点列表,提取loc和priority值,构造成目标字典列表。

解决方案二:转换为JSON后处理

利用community.general.xml_to_json过滤器将XML转为JSON,再通过JSON路径快速提取数据:

- name: Get the website 'sitemap.xml' file
  ansible.builtin.uri:
    url: "https://example.com/sitemap.xml"
    method: GET
    return_content: true
    headers:
      Accept: "application/xml"
    status_code: 200
    timeout: 5
  register: sitemap
  delegate_to: localhost

- name: Convert XML to JSON
  set_fact:
    sitemap_json: "{{ sitemap.content | community.general.xml_to_json }}"
  delegate_to: localhost

- name: Generate loc-priority dictionary list from JSON
  set_fact:
    sitemap_entries: "{{ sitemap_json.urlset.url | json_query('[].{loc: loc, priority: priority}') }}"
  delegate_to: localhost

- name: Debug the result
  ansible.builtin.debug:
    var: sitemap_entries
  delegate_to: localhost

关键说明

  • community.general.xml_to_json 将XML结构转为JSON,此时urlset下的url会自动转为列表;
  • 使用json_query过滤器直接提取每个url中的loc和priority字段,一步生成目标字典列表,代码更简洁。

内容的提问来源于stack exchange,提问作者Jaume Sabater

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 00:07:49