You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python网页爬取:导出数据至Excel遇到问题

解决方法

你遇到的问题是因为直接将BeautifulSoup返回的Tag对象集合传入pd.DataFrame(),这些对象包含完整的HTML标签结构,所以导出的Excel会出现大量无关内容。正确的做法是先把需要的笔记本名称和完整链接整理成结构化数据,再转换成DataFrame。

修正后的完整代码

import requests
from bs4 import BeautifulSoup
import pandas as pd

url = "https://webscraper.io/test-sites/e-commerce/allinone/computers/laptops"

r = requests.get(url)
soup = BeautifulSoup(r.text, "html.parser")

css_selector = {"class": "col-sm-4 col-lg-4 col-md-4"}
laptops = soup.find_all("div", attrs=css_selector)

# 初始化空列表存储结构化数据
laptop_data = []

for laptop in laptops:
    laptop_link = laptop.find('a')
    text = laptop_link.get_text(strip=True)  # 去除文本前后多余空格和换行
    href = laptop_link['href']
    full_url = f"https://webscraper.io{href}"
    # 将单条数据以字典形式存入列表
    laptop_data.append({
        "笔记本名称": text,
        "产品链接": full_url
    })

# 用结构化数据创建DataFrame并导出
df = pd.DataFrame(laptop_data)
df.to_excel("laptops_testing.xlsx", index=False, encoding="utf-8")

关键修改说明

  • 新增laptop_data列表,循环中把每台笔记本的名称和链接以{"列名": 值}的字典格式存入,确保数据是结构化的。
  • get_text(strip=True)可以清理文本中的冗余空格和换行,让导出的名称更整洁。
  • 导出时添加index=False,避免生成多余的索引列。

内容的提问来源于stack exchange,提问作者The_N00b

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 09:05:34