You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取网页特定元素并导入Pandas?菜单爬取遇属性错误求助

Fixing Your Campus Menu Scraper & Building a Pandas DataFrame

Hey Chris, let's break down how to fix those two issues you're facing—getting past that AttributeError and structuring your menu data into a clean DataFrame.

1. Fixing the AttributeError: 'NoneType' object has no attribute 'get_text'

The error pops up because you're trying to grab an <h2 class="menu-details-day"> directly under your week container, but that element is actually nested inside individual day-specific blocks. Here's the fix:

First, pull all the separate day sections from the menu, then extract the day text from each block:

# After defining your 'week' variable
day_blocks = week.find_all("div", class_="menu-details-day")

for day_block in day_blocks:
    # Now we're targeting the h2 inside each day's dedicated block
    day = day_block.find("h2", class_="menu-details-day").get_text(strip=True)
    print(day)

The strip=True cleans up extra spaces and line breaks in the output text.

2. Structuring Data into a Pandas DataFrame

Your original code returns a big unstructured text block because you're grabbing all text from the week container at once. Instead, we need to traverse the HTML hierarchy step by step—date → station → menu item—to collect each piece of data into a structured format.

Here's the full updated code that builds your desired DataFrame:

import requests
from bs4 import BeautifulSoup
import pandas as pd

# Fetch the menu page
page = requests.get("https://trinity.campusdish.com/Commerce/Catalog/Menus.aspx?LocationId=1031&PeriodId=1471&MenuDate=&Mode=week")
soup = BeautifulSoup(page.content, 'html.parser')

# Locate the main menu container
seven_day = soup.find(id="WebPartManager1_wpMenuDetails")
menu_items = seven_day.find_all(class_="menu-details")
week = menu_items[0]

# Initialize a list to store structured menu data
menu_data = []

# Iterate over each day's section
day_blocks = week.find_all("div", class_="menu-details-day")
for day_block in day_blocks:
    # Extract the full date (e.g., "Monday, October 23")
    day_date = day_block.find("h2", class_="menu-details-day").get_text(strip=True)
    
    # Iterate over each food station in the day
    stations = day_block.find_all("div", class_="menu-details-station")
    for station in stations:
        # Extract the station name (e.g., "Breakfast Bar")
        station_name = station.find("h3", class_="menu-details-station-name").get_text(strip=True)
        
        # Iterate over each menu item in the station
        items = station.find_all("div", class_="menu-details-item")
        for item in items:
            # Extract the item name (e.g., "Scrambled Eggs")
            item_name = item.get_text(strip=True)
            
            # Add the entry as a dictionary to our data list
            menu_data.append({
                "Date": day_date,
                "Station": station_name,
                "Item": item_name
            })

# Convert the structured list to a Pandas DataFrame
menu_df = pd.DataFrame(menu_data)

# Preview the result
print(menu_df.head())

How This Works:

  • We traverse the HTML in layers: first each day, then each station within the day, then each item within the station.
  • Each entry is stored as a dictionary with your three desired columns: Date, Station, Item.
  • Converting the list of dictionaries to a DataFrame is seamless with Pandas.

Quick Troubleshooting Tip:

If the selectors stop working later, run print(soup.prettify()) to view the formatted HTML of the page. This lets you check if the website updated the class names or structure of menu elements (a common occurrence!).

内容的提问来源于stack exchange,提问作者Christopher Reid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:09:49