You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

代码无报错运行后CSV文件为空,该如何排查解决?

问题排查与修复

你的代码没报错但CSV为空,核心问题是元素选择器写法错误,导致没抓取到任何内容,循环根本没执行。

错误点1:attrs参数使用错误

findAll的attrs需要传字典格式,你写的("class" == "lead session-desc")是个布尔表达式,结果为False,BeautifulSoup会用这个值去匹配元素,自然找不到任何内容。

正确写法有两种:

  • 字典形式:soup.findAll("p", attrs={"class": "lead session-desc"})
  • 更简洁的class_参数(避开Python关键字class):soup.findAll("p", class_="lead session-desc")

标题选择器同理,要改成titles = soup.findAll("h3", class_="page-title")

错误点2:多余的文件关闭语句

with是Python的上下文管理器,代码块结束后会自动关闭文件,不需要额外写f.close(),这行代码甚至会报错(f在with块外已不存在)。

修正后的完整代码

from bs4 import BeautifulSoup
import requests
import csv

page_to_scrape = requests.get("https://www.scrapethissite.com/pages/")
soup = BeautifulSoup(page_to_scrape.text, "html.parser")
# 修正选择器写法
descriptions = soup.findAll("p", class_="lead session-desc")
titles = soup.findAll("h3", class_="page-title")

with open("scrapeinformation.csv", "w", newline="", encoding="utf-8") as f:
    thewriter = csv.writer(f)
    # 写入表头让CSV更规范
    thewriter.writerow(["页面标题", "描述"])
    for title, desc in zip(titles, descriptions):
        # 清理文本里的换行和多余空格
        clean_title = title.text.strip()
        clean_desc = desc.text.strip()
        print(f"{clean_title} - {clean_desc}")
        thewriter.writerow([clean_title, clean_desc])

额外优化建议

  • 加上encoding="utf-8",避免CSV出现中文乱码
  • 用strip()清理文本换行和冗余空格,让输出更整洁
  • 写入表头,让CSV结构更清晰

内容的提问来源于stack exchange,提问作者Camol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 11:42:05