如何提取网页特定元素并导入Pandas?菜单爬取遇属性错误求助
Hey Chris, let's break down how to fix those two issues you're facing—getting past that AttributeError and structuring your menu data into a clean DataFrame.
1. Fixing the AttributeError: 'NoneType' object has no attribute 'get_text'
The error pops up because you're trying to grab an <h2 class="menu-details-day"> directly under your week container, but that element is actually nested inside individual day-specific blocks. Here's the fix:
First, pull all the separate day sections from the menu, then extract the day text from each block:
# After defining your 'week' variable day_blocks = week.find_all("div", class_="menu-details-day") for day_block in day_blocks: # Now we're targeting the h2 inside each day's dedicated block day = day_block.find("h2", class_="menu-details-day").get_text(strip=True) print(day)
The strip=True cleans up extra spaces and line breaks in the output text.
2. Structuring Data into a Pandas DataFrame
Your original code returns a big unstructured text block because you're grabbing all text from the week container at once. Instead, we need to traverse the HTML hierarchy step by step—date → station → menu item—to collect each piece of data into a structured format.
Here's the full updated code that builds your desired DataFrame:
import requests from bs4 import BeautifulSoup import pandas as pd # Fetch the menu page page = requests.get("https://trinity.campusdish.com/Commerce/Catalog/Menus.aspx?LocationId=1031&PeriodId=1471&MenuDate=&Mode=week") soup = BeautifulSoup(page.content, 'html.parser') # Locate the main menu container seven_day = soup.find(id="WebPartManager1_wpMenuDetails") menu_items = seven_day.find_all(class_="menu-details") week = menu_items[0] # Initialize a list to store structured menu data menu_data = [] # Iterate over each day's section day_blocks = week.find_all("div", class_="menu-details-day") for day_block in day_blocks: # Extract the full date (e.g., "Monday, October 23") day_date = day_block.find("h2", class_="menu-details-day").get_text(strip=True) # Iterate over each food station in the day stations = day_block.find_all("div", class_="menu-details-station") for station in stations: # Extract the station name (e.g., "Breakfast Bar") station_name = station.find("h3", class_="menu-details-station-name").get_text(strip=True) # Iterate over each menu item in the station items = station.find_all("div", class_="menu-details-item") for item in items: # Extract the item name (e.g., "Scrambled Eggs") item_name = item.get_text(strip=True) # Add the entry as a dictionary to our data list menu_data.append({ "Date": day_date, "Station": station_name, "Item": item_name }) # Convert the structured list to a Pandas DataFrame menu_df = pd.DataFrame(menu_data) # Preview the result print(menu_df.head())
How This Works:
- We traverse the HTML in layers: first each day, then each station within the day, then each item within the station.
- Each entry is stored as a dictionary with your three desired columns:
Date,Station,Item. - Converting the list of dictionaries to a DataFrame is seamless with Pandas.
Quick Troubleshooting Tip:
If the selectors stop working later, run print(soup.prettify()) to view the formatted HTML of the page. This lets you check if the website updated the class names or structure of menu elements (a common occurrence!).
内容的提问来源于stack exchange,提问作者Christopher Reid

