You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从指定网页提取课程信息并构建数据集?

Solution

Here's how to extend your code to extract the required data and format it into a structured dataset:

Step-by-Step Explanation & Code

First, we'll use BeautifulSoup to navigate the page's HTML structure, then extract the program name and course details per semester. We'll use pandas to organize the data into a usable dataset (install it with pip install pandas if you haven't already).

from bs4 import BeautifulSoup
from urllib.request import urlopen
import pandas as pd

url = "https://www.lusem.lu.se/study/bachelors/economy-and-society/programme-structure"
page = urlopen(url)
html = page.read().decode("utf-8")
soup = BeautifulSoup(html, "html.parser")

# Extract the main program name
program_name = soup.find("h1").get_text(strip=True)

# Initialize a list to store course data
course_records = []

# Iterate over each semester section
semester_headings = soup.find_all("h2", string=lambda text: "Semester" in text.strip() if text else False)
for heading in semester_headings:
    semester = heading.get_text(strip=True)
    # Get the table containing courses for this semester
    course_table = heading.find_next("table")
    if not course_table:
        continue
    
    # Skip the header row (first row in the table)
    course_rows = course_table.find_all("tr")[1:]
    for row in course_rows:
        columns = row.find_all("td")
        if len(columns) < 2:
            continue  # Skip rows with incomplete data
        
        course_code = columns[0].get_text(strip=True)
        course_name = columns[1].get_text(strip=True)
        
        # Add the course to our records
        course_records.append({
            "Program Name": program_name,
            "Semester": semester,
            "Course Code": course_code,
            "Course Name": course_name
        })

# Convert to a structured DataFrame
course_dataset = pd.DataFrame(course_records)

# Optional: Save to CSV for easy access
course_dataset.to_csv("economy_and_society_program_courses.csv", index=False, encoding="utf-8")

# Print results
print(f"Program Name: {program_name}\n")
print("Course Dataset:")
print(course_dataset)

Key Details:

  • Program Name Extraction: The main program title lives in the page's top <h1> tag, so we target that directly.
  • Semester Navigation: We filter all <h2> tags containing "Semester" to isolate each semester's section.
  • Course Table Parsing: Each semester's courses are stored in a table directly below the semester heading. We skip the header row (with column labels) and pull the course code (first column) and name (second column) from each data row.
  • Structured Dataset: Using pandas.DataFrame converts raw extracted data into a clean, tabular format that can be exported to CSV or used for further analysis.

Notes:

  • If the page structure changes (e.g., courses move from tables to lists), adjust the code to target <ul>/<li> elements instead of tables.
  • The code skips rows with incomplete data to avoid runtime errors.

内容的提问来源于stack exchange,提问作者Neg K

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 12:26:14