You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

含特殊字符的科目名称遍历方法及BeautifulSoup匹配属性错误的代码修复方案

Fixing MSBTE Result Scraper: Handling Special Characters & Robust Subject Traversal

Let's tackle the two main issues with your scraper: matching subjects with special characters and building a flexible traversal logic that adapts to subject changes. Here's a refined solution with clear explanations:

Key Issues Addressed

  1. Special Character Matching: Exact text matching fails when subjects have extra spaces, special symbols, or formatting differences. We'll use fuzzy matching with cleaned text to avoid this.
  2. Hardcoded Indexes: Replaced manual find_next loops with row-based extraction—grabbing entire table rows for each subject makes the code resilient to changes in subject structure.
  3. Error Resilience: Added checks to avoid AttributeError when subjects aren't found, plus optimized file operations for better performance.

Refactored Code

import requests
from bs4 import BeautifulSoup
import csv
import re

def clean_text(text):
    """Clean text by stripping whitespace and normalizing special characters"""
    if not text:
        return ""
    # Remove extra spaces, newlines, and normalize special characters
    return re.sub(r'\s+', ' ', text.strip()).upper()

# Base URL template (avoids hardcoding dynamic parts)
BASE_URL = "https://msbte.org.in/DISRESLIVE2021CRSLDSEP/COV6139QS21LIVEResult/SeatNumber/30/{seat_number}Marksheet.html"
# Subject list (we'll match these case-insensitively with cleaned text)
SUBJECTS = [
    "MANAGEMENT",
    "PROGRAMMING WITH PYTHON",
    "MOBILE APPLICATION DEVELOPMENT",
    "EMERGING TRENDS IN COMPUTER AND INFORMATION TECHNOLOGY",  # Fixed typo
    "NETWORK AND INFORMATION SECURITY",
    "ENTREPRENEURSHIP DEVELOPMENT",  # Fixed typo
    "CAPSTONE PROJECT EXECUTION & REPORT WRITING"
]
START_SEAT = 302060
END_SEAT = 302065

# Open CSV file once (more efficient than opening/closing per seat)
with open("C:\\sem6.csv", 'w', newline='', encoding='utf-8') as f:
    writer = csv.writer(f)
    # Build header row for clarity
    header = ["Seat Number", "Student Name"]
    for sub in SUBJECTS:
        # Add columns based on subject score count
        if sub in ["PROGRAMMING WITH PYTHON", "MOBILE APPLICATION DEVELOPMENT", "NETWORK AND INFORMATION SECURITY"]:
            header.extend([f"{sub} - Score 1", f"{sub} - Score 2", f"{sub} - Score 3", f"{sub} - Score 4"])
        else:
            header.extend([f"{sub} - Score 1", f"{sub} - Score 2"])
    header.extend(["Total", "Percentage", "Remarks"])
    writer.writerow(header)

    for seat_number in range(START_SEAT, END_SEAT):
        marks = []
        url = BASE_URL.format(seat_number=seat_number)
        try:
            r = requests.get(url, timeout=10)
            r.raise_for_status()  # Catch HTTP errors
            soup = BeautifulSoup(r.content, 'html.parser')

            # Extract student basic details
            student_table = soup.find_all('table')[0]
            name = student_table.find_all('tr')[0].find_all('td')[1].text.strip()
            seat_no = soup.find('td', text=lambda t: t and "SEAT NO." in t.strip()).find_next('td').text.strip()
            marks.append(seat_no)
            marks.append(name)

            # Extract subject marks from the main table
            subject_table = soup.find_all('table')[1]
            subject_rows = subject_table.find_all('tr')[2:]  # Skip header rows

            for target_subject in SUBJECTS:
                cleaned_target = clean_text(target_subject)
                subject_found = False
                for row in subject_rows:
                    cells = row.find_all('td')
                    if not cells:
                        continue
                    # Clean the subject name from the current row
                    row_subject = clean_text(cells[0].text)
                    # Fuzzy match: check if either text contains the other
                    if cleaned_target in row_subject or row_subject in cleaned_target:
                        subject_found = True
                        # Extract relevant scores based on subject type
                        if target_subject in ["PROGRAMMING WITH PYTHON", "MOBILE APPLICATION DEVELOPMENT", "NETWORK AND INFORMATION SECURITY"]:
                            score_cols = [cells[5].text.strip(), cells[14].text.strip(), cells[23].text.strip(), cells[32].text.strip()]
                        else:
                            score_cols = [cells[5].text.strip(), cells[14].text.strip()]
                        marks.extend(score_cols)
                        break
                # Add empty values if subject isn't found to maintain CSV structure
                if not subject_found:
                    if target_subject in ["PROGRAMMING WITH PYTHON", "MOBILE APPLICATION DEVELOPMENT", "NETWORK AND INFORMATION SECURITY"]:
                        marks.extend(["", "", "", ""])
                    else:
                        marks.extend(["", ""])

            # Extract total, percentage, remarks
            result_summary = soup.find_all('table')[2]
            total = result_summary.find_all('tr')[1].find_all('td')[2].text.strip()
            percentage = result_summary.find_all('tr')[1].find_all('td')[1].text.strip()
            remarks = result_summary.find_all('tr')[2].find_all('td')[1].text.strip()
            marks.extend([total, percentage, remarks])

            # Write the row to CSV
            writer.writerow(marks)
            print(f"Successfully scraped seat number: {seat_number}")

        except Exception as e:
            print(f"Error scraping seat {seat_number}: {str(e)}")
            # Track failures with a placeholder row
            writer.writerow([seat_number, "Scraping Error"] + [""]*(len(header)-2))

print("Scraping completed!")

What's Improved?

  • clean_text Function: Normalizes text to eliminate formatting differences (extra spaces, newlines) that break exact matches.
  • Fuzzy Matching: Uses a lambda function to check if the cleaned target subject exists in the row's text, handling special characters and minor typos.
  • Row-Based Extraction: Iterates over table rows instead of jumping through cells, making the code more reliable if the page structure changes.
  • Error Handling: Try/except blocks catch HTTP errors and missing elements, ensuring the scraper runs to completion even if some seats fail.
  • CSV Header: Adds a descriptive header to make the output CSV easier to analyze.
  • Optimized File I/O: Opens the CSV file once instead of per seat, reducing overhead and improving speed.

内容的提问来源于stack exchange,提问作者user6903964

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 15:22:47