使用BeautifulSoup和Requests爬取LMS日历报错:NoneType不可下标
Hey there! Let's break down why that TypeError: 'NoneType' object is not subscriptable is popping up and get your script working smoothly.
First, the Root Cause of the Error
That error on line 5 (first_table = soup.find("table")[1]) happens because soup.find("table") is returning None—meaning BeautifulSoup couldn't find any <table> element matching your query. When you try to index a None value with [1], Python throws that error.
There are two main reasons this is happening:
- You're not logged into the LMS: The URL
https://gatech.instructure.com/calendarrequires authentication. When you userequests.get()without logging in, you're getting the login page HTML instead of the actual calendar page—so there's no table with the structure you're looking for. - Incorrect use of
find()vsfind_all():find()returns only the first matching element (orNoneif nothing is found). If you're trying to access the second table on the page, you need to usefind_all()(which returns a list of all matches) and then index into that list.
Step-by-Step Fixes
1. Handle Authentication (Critical!)
Canvas (Georgia Tech's LMS) requires you to be logged in to access the calendar. Here's a simple way to do this using a session and cookies (you can grab your cookies from your browser's dev tools under Application > Cookies):
import requests from bs4 import BeautifulSoup # Create a session to persist login state session = requests.Session() # Add your Canvas cookies here (copy values from your browser) cookies = { '_canvas_session': 'YOUR_CANVAS_SESSION_COOKIE', # Add any other required cookies listed for the domain } # Update the session with your cookies session.cookies.update(cookies) # Now access the calendar page response = session.get('https://gatech.instructure.com/calendar') response.raise_for_status() # Throws an error if the request fails (e.g., invalid cookies) website = response.text soup = BeautifulSoup(website, 'lxml')
2. Fix the find()/find_all() Mistakes
Let's correct the lines where you're using indexing on find() results (which returns a single element, not a list):
- Line 5: Replace
soup.find("table")[1]withsoup.find_all("table")[1](sincefind_all()returns a list, and we want the second table). Add a check to make sure the list has enough elements first. - Line 8:
big_body.find('div', ...)[x]should bebig_body.find_all('div', {'class': 'fc-row fc-week fc-widget-content'})[x] - Line 10:
week_table.find('td', ...)[2]is incorrect—find()returns a single<td>element, so indexing with[2]doesn't make sense. We'll add a check to ensure the td exists instead. - Line 11: Use
find_all()instead offind()to get all assignment links, not just the first one.
Revised Working Code (With Error Checks)
Here's the updated script with safeguards to prevent NoneType errors and correct usage:
import requests from bs4 import BeautifulSoup # Set up session with authentication cookies session = requests.Session() cookies = { '_canvas_session': 'YOUR_SESSION_COOKIE_HERE', # Add other necessary cookies from your browser } session.cookies.update(cookies) # Fetch calendar page try: response = session.get('https://gatech.instructure.com/calendar') response.raise_for_status() except requests.exceptions.RequestException as e: print(f"Failed to access the calendar: {e}") exit() soup = BeautifulSoup(response.text, 'lxml') # Get the second table (index 1) tables = soup.find_all("table") if len(tables) < 2: print("Couldn't find the expected table on the page.") exit() first_table = tables[1] # Get the calendar body big_body = first_table.find("tbody", {'class': 'fc-body'}) if not big_body: print("Couldn't find the calendar body.") exit() # Get user input for week number try: x = int(input("What week of the month? (0 for first week) ")) except ValueError: print("Please enter a valid number.") exit() # Get the specified week row week_rows = big_body.find_all('div', {'class': 'fc-row fc-week fc-widget-content'}) if x >= len(week_rows): print(f"Only {len(week_rows)} weeks found. Please enter a smaller number.") exit() week_div = week_rows[x] # Get the week's table week_tables = week_div.find_all('table') if len(week_tables) < 2: print("Couldn't find the week table.") exit() week_table = week_tables[1] # Get the current day (highlighted cell) current_day = week_table.find('td', attrs={'bgcolor': '#f5f5f5'}) if not current_day: print("Couldn't find the current day's cell.") exit() # Get all assignment links assignments = current_day.find_all('a', {'class': 'fc-day-grid-event'}) if not assignments: print("No assignments found for today.") exit() # Collect assignment titles data = [] for assignment in assignments: title = assignment.get('title') if title: data.append(title) print("Today's Assignments:") for item in data: print(f"- {item}")
Quick Tips for Future Scraping
- Always add checks to make sure elements exist before trying to access them (this prevents most
NoneTypeerrors). - Use
find_all()when you need multiple elements; reservefind()for when you want a single match. - Don't skip authentication—most LMS pages require logged-in access to load content.
内容的提问来源于stack exchange,提问作者Scott Janssen

