You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup和Requests爬取LMS日历报错:NoneType不可下标

Fixing the TypeError in Your LMS Calendar Scraper

Hey there! Let's break down why that TypeError: 'NoneType' object is not subscriptable is popping up and get your script working smoothly.

First, the Root Cause of the Error

That error on line 5 (first_table = soup.find("table")[1]) happens because soup.find("table") is returning None—meaning BeautifulSoup couldn't find any <table> element matching your query. When you try to index a None value with [1], Python throws that error.

There are two main reasons this is happening:

  1. You're not logged into the LMS: The URL https://gatech.instructure.com/calendar requires authentication. When you use requests.get() without logging in, you're getting the login page HTML instead of the actual calendar page—so there's no table with the structure you're looking for.
  2. Incorrect use of find() vs find_all(): find() returns only the first matching element (or None if nothing is found). If you're trying to access the second table on the page, you need to use find_all() (which returns a list of all matches) and then index into that list.

Step-by-Step Fixes

1. Handle Authentication (Critical!)

Canvas (Georgia Tech's LMS) requires you to be logged in to access the calendar. Here's a simple way to do this using a session and cookies (you can grab your cookies from your browser's dev tools under Application > Cookies):

import requests
from bs4 import BeautifulSoup

# Create a session to persist login state
session = requests.Session()

# Add your Canvas cookies here (copy values from your browser)
cookies = {
    '_canvas_session': 'YOUR_CANVAS_SESSION_COOKIE',
    # Add any other required cookies listed for the domain
}

# Update the session with your cookies
session.cookies.update(cookies)

# Now access the calendar page
response = session.get('https://gatech.instructure.com/calendar')
response.raise_for_status()  # Throws an error if the request fails (e.g., invalid cookies)
website = response.text
soup = BeautifulSoup(website, 'lxml')

2. Fix the find()/find_all() Mistakes

Let's correct the lines where you're using indexing on find() results (which returns a single element, not a list):

  • Line 5: Replace soup.find("table")[1] with soup.find_all("table")[1] (since find_all() returns a list, and we want the second table). Add a check to make sure the list has enough elements first.
  • Line 8: big_body.find('div', ...)[x] should be big_body.find_all('div', {'class': 'fc-row fc-week fc-widget-content'})[x]
  • Line 10: week_table.find('td', ...)[2] is incorrect—find() returns a single <td> element, so indexing with [2] doesn't make sense. We'll add a check to ensure the td exists instead.
  • Line 11: Use find_all() instead of find() to get all assignment links, not just the first one.

Revised Working Code (With Error Checks)

Here's the updated script with safeguards to prevent NoneType errors and correct usage:

import requests
from bs4 import BeautifulSoup

# Set up session with authentication cookies
session = requests.Session()
cookies = {
    '_canvas_session': 'YOUR_SESSION_COOKIE_HERE',
    # Add other necessary cookies from your browser
}
session.cookies.update(cookies)

# Fetch calendar page
try:
    response = session.get('https://gatech.instructure.com/calendar')
    response.raise_for_status()
except requests.exceptions.RequestException as e:
    print(f"Failed to access the calendar: {e}")
    exit()

soup = BeautifulSoup(response.text, 'lxml')

# Get the second table (index 1)
tables = soup.find_all("table")
if len(tables) < 2:
    print("Couldn't find the expected table on the page.")
    exit()
first_table = tables[1]

# Get the calendar body
big_body = first_table.find("tbody", {'class': 'fc-body'})
if not big_body:
    print("Couldn't find the calendar body.")
    exit()

# Get user input for week number
try:
    x = int(input("What week of the month? (0 for first week) "))
except ValueError:
    print("Please enter a valid number.")
    exit()

# Get the specified week row
week_rows = big_body.find_all('div', {'class': 'fc-row fc-week fc-widget-content'})
if x >= len(week_rows):
    print(f"Only {len(week_rows)} weeks found. Please enter a smaller number.")
    exit()
week_div = week_rows[x]

# Get the week's table
week_tables = week_div.find_all('table')
if len(week_tables) < 2:
    print("Couldn't find the week table.")
    exit()
week_table = week_tables[1]

# Get the current day (highlighted cell)
current_day = week_table.find('td', attrs={'bgcolor': '#f5f5f5'})
if not current_day:
    print("Couldn't find the current day's cell.")
    exit()

# Get all assignment links
assignments = current_day.find_all('a', {'class': 'fc-day-grid-event'})
if not assignments:
    print("No assignments found for today.")
    exit()

# Collect assignment titles
data = []
for assignment in assignments:
    title = assignment.get('title')
    if title:
        data.append(title)

print("Today's Assignments:")
for item in data:
    print(f"- {item}")

Quick Tips for Future Scraping

  • Always add checks to make sure elements exist before trying to access them (this prevents most NoneType errors).
  • Use find_all() when you need multiple elements; reserve find() for when you want a single match.
  • Don't skip authentication—most LMS pages require logged-in access to load content.

内容的提问来源于stack exchange,提问作者Scott Janssen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 15:47:28