如何用Python从指定网页提取课程信息并构建数据集?
Solution
Here's how to extend your code to extract the required data and format it into a structured dataset:
Step-by-Step Explanation & Code
First, we'll use BeautifulSoup to navigate the page's HTML structure, then extract the program name and course details per semester. We'll use pandas to organize the data into a usable dataset (install it with pip install pandas if you haven't already).
from bs4 import BeautifulSoup from urllib.request import urlopen import pandas as pd url = "https://www.lusem.lu.se/study/bachelors/economy-and-society/programme-structure" page = urlopen(url) html = page.read().decode("utf-8") soup = BeautifulSoup(html, "html.parser") # Extract the main program name program_name = soup.find("h1").get_text(strip=True) # Initialize a list to store course data course_records = [] # Iterate over each semester section semester_headings = soup.find_all("h2", string=lambda text: "Semester" in text.strip() if text else False) for heading in semester_headings: semester = heading.get_text(strip=True) # Get the table containing courses for this semester course_table = heading.find_next("table") if not course_table: continue # Skip the header row (first row in the table) course_rows = course_table.find_all("tr")[1:] for row in course_rows: columns = row.find_all("td") if len(columns) < 2: continue # Skip rows with incomplete data course_code = columns[0].get_text(strip=True) course_name = columns[1].get_text(strip=True) # Add the course to our records course_records.append({ "Program Name": program_name, "Semester": semester, "Course Code": course_code, "Course Name": course_name }) # Convert to a structured DataFrame course_dataset = pd.DataFrame(course_records) # Optional: Save to CSV for easy access course_dataset.to_csv("economy_and_society_program_courses.csv", index=False, encoding="utf-8") # Print results print(f"Program Name: {program_name}\n") print("Course Dataset:") print(course_dataset)
Key Details:
- Program Name Extraction: The main program title lives in the page's top
<h1>tag, so we target that directly. - Semester Navigation: We filter all
<h2>tags containing "Semester" to isolate each semester's section. - Course Table Parsing: Each semester's courses are stored in a table directly below the semester heading. We skip the header row (with column labels) and pull the course code (first column) and name (second column) from each data row.
- Structured Dataset: Using
pandas.DataFrameconverts raw extracted data into a clean, tabular format that can be exported to CSV or used for further analysis.
Notes:
- If the page structure changes (e.g., courses move from tables to lists), adjust the code to target
<ul>/<li>elements instead of tables. - The code skips rows with incomplete data to avoid runtime errors.
内容的提问来源于stack exchange,提问作者Neg K
相关产品推荐
相关产品推荐

