You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup获取XML站点地图链接遇问题求助

Let's fix your sitemap scraping and Excel export issues step by step:

1. Resolve the Parser Error

First, let's tackle that FeatureNotFound error head-on:

  • Filename conflict: Your script is named lxml.py, which clashes with the official lxml library. Rename your script to something like sitemap_scraper.py immediately—this is preventing Python from importing the actual lxml parser correctly.
  • Incorrect parser name: "lxml.xml" isn't a valid parser identifier. Use "lxml" or "xml" instead.
  • Undefined variable: You used site in soup(site, "lxml.xml") but never defined it—this should be page_html, the raw content you fetched from the sitemap.

2. Fix Sitemap Parsing Logic

Sitemaps are XML files, not HTML—there are no <table> elements here. You need to target the <url> tags and extract their child <loc> elements, which contain the actual URLs you want.

3. Add Excel Export Functionality

We'll use pandas to easily collect and export the URLs to an Excel file—it's straightforward and handles formatting automatically.

Full Working Code

import bs4
from urllib.request import urlopen as uReq
from bs4 import BeautifulSoup as soup
import pandas as pd

# Fetch the sitemap content
my_url = 'https://www.example.com/sitemap.xml'
uClient = uReq(my_url)
page_html = uClient.read()
uClient.close()

# Parse the XML sitemap
page_soup = soup(page_html, "xml")  # Use valid XML parser (requires lxml)

# Extract all URL locations from the sitemap
url_elements = page_soup.findAll("loc")
urls = [element.text.strip() for element in url_elements]

# Export collected URLs to Excel
df = pd.DataFrame({"Sitemap URLs": urls})
df.to_excel("sitemap_urls.xlsx", index=False)

print(f"Successfully exported {len(urls)} URLs to sitemap_urls.xlsx")

Setup Notes

  1. Install the required packages if you haven't already:
    pip install lxml pandas openpyxl
    
  2. Double-check your script name doesn't match any library names (like lxml.py or pandas.py) to avoid import conflicts.

This code will fetch all URLs from your sitemap, store them in a structured DataFrame, and export them to a clean, ready-to-use Excel file.

内容的提问来源于stack exchange,提问作者noussama

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:18:45