含特殊字符的科目名称遍历方法及BeautifulSoup匹配属性错误的代码修复方案
Fixing MSBTE Result Scraper: Handling Special Characters & Robust Subject Traversal
Let's tackle the two main issues with your scraper: matching subjects with special characters and building a flexible traversal logic that adapts to subject changes. Here's a refined solution with clear explanations:
Key Issues Addressed
- Special Character Matching: Exact text matching fails when subjects have extra spaces, special symbols, or formatting differences. We'll use fuzzy matching with cleaned text to avoid this.
- Hardcoded Indexes: Replaced manual
find_nextloops with row-based extraction—grabbing entire table rows for each subject makes the code resilient to changes in subject structure. - Error Resilience: Added checks to avoid
AttributeErrorwhen subjects aren't found, plus optimized file operations for better performance.
Refactored Code
import requests from bs4 import BeautifulSoup import csv import re def clean_text(text): """Clean text by stripping whitespace and normalizing special characters""" if not text: return "" # Remove extra spaces, newlines, and normalize special characters return re.sub(r'\s+', ' ', text.strip()).upper() # Base URL template (avoids hardcoding dynamic parts) BASE_URL = "https://msbte.org.in/DISRESLIVE2021CRSLDSEP/COV6139QS21LIVEResult/SeatNumber/30/{seat_number}Marksheet.html" # Subject list (we'll match these case-insensitively with cleaned text) SUBJECTS = [ "MANAGEMENT", "PROGRAMMING WITH PYTHON", "MOBILE APPLICATION DEVELOPMENT", "EMERGING TRENDS IN COMPUTER AND INFORMATION TECHNOLOGY", # Fixed typo "NETWORK AND INFORMATION SECURITY", "ENTREPRENEURSHIP DEVELOPMENT", # Fixed typo "CAPSTONE PROJECT EXECUTION & REPORT WRITING" ] START_SEAT = 302060 END_SEAT = 302065 # Open CSV file once (more efficient than opening/closing per seat) with open("C:\\sem6.csv", 'w', newline='', encoding='utf-8') as f: writer = csv.writer(f) # Build header row for clarity header = ["Seat Number", "Student Name"] for sub in SUBJECTS: # Add columns based on subject score count if sub in ["PROGRAMMING WITH PYTHON", "MOBILE APPLICATION DEVELOPMENT", "NETWORK AND INFORMATION SECURITY"]: header.extend([f"{sub} - Score 1", f"{sub} - Score 2", f"{sub} - Score 3", f"{sub} - Score 4"]) else: header.extend([f"{sub} - Score 1", f"{sub} - Score 2"]) header.extend(["Total", "Percentage", "Remarks"]) writer.writerow(header) for seat_number in range(START_SEAT, END_SEAT): marks = [] url = BASE_URL.format(seat_number=seat_number) try: r = requests.get(url, timeout=10) r.raise_for_status() # Catch HTTP errors soup = BeautifulSoup(r.content, 'html.parser') # Extract student basic details student_table = soup.find_all('table')[0] name = student_table.find_all('tr')[0].find_all('td')[1].text.strip() seat_no = soup.find('td', text=lambda t: t and "SEAT NO." in t.strip()).find_next('td').text.strip() marks.append(seat_no) marks.append(name) # Extract subject marks from the main table subject_table = soup.find_all('table')[1] subject_rows = subject_table.find_all('tr')[2:] # Skip header rows for target_subject in SUBJECTS: cleaned_target = clean_text(target_subject) subject_found = False for row in subject_rows: cells = row.find_all('td') if not cells: continue # Clean the subject name from the current row row_subject = clean_text(cells[0].text) # Fuzzy match: check if either text contains the other if cleaned_target in row_subject or row_subject in cleaned_target: subject_found = True # Extract relevant scores based on subject type if target_subject in ["PROGRAMMING WITH PYTHON", "MOBILE APPLICATION DEVELOPMENT", "NETWORK AND INFORMATION SECURITY"]: score_cols = [cells[5].text.strip(), cells[14].text.strip(), cells[23].text.strip(), cells[32].text.strip()] else: score_cols = [cells[5].text.strip(), cells[14].text.strip()] marks.extend(score_cols) break # Add empty values if subject isn't found to maintain CSV structure if not subject_found: if target_subject in ["PROGRAMMING WITH PYTHON", "MOBILE APPLICATION DEVELOPMENT", "NETWORK AND INFORMATION SECURITY"]: marks.extend(["", "", "", ""]) else: marks.extend(["", ""]) # Extract total, percentage, remarks result_summary = soup.find_all('table')[2] total = result_summary.find_all('tr')[1].find_all('td')[2].text.strip() percentage = result_summary.find_all('tr')[1].find_all('td')[1].text.strip() remarks = result_summary.find_all('tr')[2].find_all('td')[1].text.strip() marks.extend([total, percentage, remarks]) # Write the row to CSV writer.writerow(marks) print(f"Successfully scraped seat number: {seat_number}") except Exception as e: print(f"Error scraping seat {seat_number}: {str(e)}") # Track failures with a placeholder row writer.writerow([seat_number, "Scraping Error"] + [""]*(len(header)-2)) print("Scraping completed!")
What's Improved?
clean_textFunction: Normalizes text to eliminate formatting differences (extra spaces, newlines) that break exact matches.- Fuzzy Matching: Uses a lambda function to check if the cleaned target subject exists in the row's text, handling special characters and minor typos.
- Row-Based Extraction: Iterates over table rows instead of jumping through cells, making the code more reliable if the page structure changes.
- Error Handling: Try/except blocks catch HTTP errors and missing elements, ensuring the scraper runs to completion even if some seats fail.
- CSV Header: Adds a descriptive header to make the output CSV easier to analyze.
- Optimized File I/O: Opens the CSV file once instead of per seat, reducing overhead and improving speed.
内容的提问来源于stack exchange,提问作者user6903964
相关产品推荐
相关产品推荐

