You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests与BeautifulSoup爬取目标内容报错及CSV导出问题

解决爬取Code.org课程目标的错误及生成CSV方案

错误原因分析

AttributeError: 'NoneType' object has no attribute 'next_sibling' 是因为 soup.find(text="Students will be able to:") 没有匹配到任何文本节点,返回了None。问题出在:

  • 页面中的目标文本被嵌套在HTML标签内,并非独立文本节点
  • 文本可能包含空格、换行符,导致精确匹配失效

修正后的单课程爬取代码

改用模糊匹配定位目标文本,再获取对应的学习目标内容:

import requests
from bs4 import BeautifulSoup

url = "https://studio.code.org/s/web-development-2023/lessons/1"
response = requests.get(url)
soup = BeautifulSoup(response.text, "html.parser")

# 模糊匹配包含目标文本的节点
target_text = soup.find(string=lambda t: t and "Students will be able to:" in t)
if target_text:
    # 找到目标文本所在的容器,提取列表形式的目标内容
    objectives_container = target_text.find_parent("div", class_="content")
    objectives = [li.get_text(strip=True) for li in objectives_container.find_all("li")]
    print("\n".join(objectives))
else:
    print("未找到学习目标内容")

多课程CSV生成方案

遍历课程ID,收集每个课程的名称和学习目标,批量写入CSV文件:

import requests
from bs4 import BeautifulSoup
import csv
import time

# 定义要爬取的课程范围,示例为第1到第5课
lesson_range = range(1, 6)
csv_rows = []

for lesson_id in lesson_range:
    url = f"https://studio.code.org/s/web-development-2023/lessons/{lesson_id}"
    response = requests.get(url)
    soup = BeautifulSoup(response.text, "html.parser")
    
    # 提取课程标题
    lesson_title = soup.find("h1", class_="lesson-title").get_text(strip=True)
    
    # 提取学习目标
    target_text = soup.find(string=lambda t: t and "Students will be able to:" in t)
    objectives = []
    if target_text:
        objectives_container = target_text.find_parent("div", class_="content")
        objectives = [li.get_text(strip=True) for li in objectives_container.find_all("li")]
    
    # 整理为CSV行数据
    csv_rows.append({
        "课程ID": lesson_id,
        "课程标题": lesson_title,
        "学习目标": "; ".join(objectives)
    })
    
    # 控制请求间隔,避免触发反爬限制
    time.sleep(1)

# 写入CSV文件(utf-8-sig确保中文正常显示)
with open("课程学习目标.csv", "w", encoding="utf-8-sig", newline="") as f:
    writer = csv.DictWriter(f, fieldnames=["课程ID", "课程标题", "学习目标"])
    writer.writeheader()
    writer.writerows(csv_rows)

print("CSV文件已生成完成")

注意事项

  • 若网站页面结构更新,需重新检查HTML标签和class属性,调整定位逻辑
  • 可根据实际需求修改lesson_range的范围,爬取更多课程

内容的提问来源于stack exchange,提问作者SK Kim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 15:40:56