You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取导出Excel问题:如何提取span标签内纯文本内容

问题解决:提取span标签内的纯文本内容

问题原因

你当前代码中直接将BeautifulSoup找到的span标签对象添加到列表中,导出时自然会输出完整的HTML标签代码。核心是要提取标签内部的纯文本内容。

修改后的代码

import pandas
import requests
from bs4 import BeautifulSoup

webpage = requests.get("https://www.lego.com/nl-nl/pick-and-build/pick-a-brick?page=1&perPage=400&filters.i0.key=variants.attributes.colourId&filters.i0.values.i0=323")
content = webpage.content
result = BeautifulSoup(content, 'html.parser')
products = result.find_all("span", {"class": "ElementLeaf_elementTitle__xda82"})

names = []
for item in products:
    # 提取标签内的纯文本,替换原有的names.append(item)
    names.append(item.text.strip())  # strip()去除文本前后的空白字符,让结果更整洁

data = list(zip(names))
d = pandas.DataFrame(data, columns=['Name'])

try:   
    d.to_excel(r"C:\Users\minib\PycharmProjects\pythonProject\scraper\lego.xlsx")  # 加r避免路径反斜杠被当作转义字符
except:
    print("Something went wrong! Please check your code.")
else:
    print("Web data successfully written to Excel.")
finally:
    print("Quitting the program. Bye!")

关键修改点

  • 替换names.append(item)为names.append(item.text.strip()):item.text获取标签包裹的纯文本内容,strip()可以清除文本前后的换行、空格等冗余字符。
  • 路径前添加r:避免Windows路径中的反斜杠被Python当作转义字符处理,防止路径解析报错。

内容的提问来源于stack exchange,提问作者minibral

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 12:45:11