代码无报错运行后CSV文件为空,该如何排查解决?
问题排查与修复
你的代码没报错但CSV为空,核心问题是元素选择器写法错误,导致没抓取到任何内容,循环根本没执行。
错误点1:attrs参数使用错误
findAll的attrs需要传字典格式,你写的("class" == "lead session-desc")是个布尔表达式,结果为False,BeautifulSoup会用这个值去匹配元素,自然找不到任何内容。
正确写法有两种:
- 字典形式:
soup.findAll("p", attrs={"class": "lead session-desc"}) - 更简洁的
class_参数(避开Python关键字class):soup.findAll("p", class_="lead session-desc")
标题选择器同理,要改成titles = soup.findAll("h3", class_="page-title")
错误点2:多余的文件关闭语句
with是Python的上下文管理器,代码块结束后会自动关闭文件,不需要额外写f.close(),这行代码甚至会报错(f在with块外已不存在)。
修正后的完整代码
from bs4 import BeautifulSoup import requests import csv page_to_scrape = requests.get("https://www.scrapethissite.com/pages/") soup = BeautifulSoup(page_to_scrape.text, "html.parser") # 修正选择器写法 descriptions = soup.findAll("p", class_="lead session-desc") titles = soup.findAll("h3", class_="page-title") with open("scrapeinformation.csv", "w", newline="", encoding="utf-8") as f: thewriter = csv.writer(f) # 写入表头让CSV更规范 thewriter.writerow(["页面标题", "描述"]) for title, desc in zip(titles, descriptions): # 清理文本里的换行和多余空格 clean_title = title.text.strip() clean_desc = desc.text.strip() print(f"{clean_title} - {clean_desc}") thewriter.writerow([clean_title, clean_desc])
额外优化建议
- 加上
encoding="utf-8",避免CSV出现中文乱码 - 用
strip()清理文本换行和冗余空格,让输出更整洁 - 写入表头,让CSV结构更清晰
内容的提问来源于stack exchange,提问作者Camol
相关产品推荐
相关产品推荐

