You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

大数量级MCQ考试特殊计分系统实现优化问询

Optimizing Your Competitive Exam Scoring Code for Large Datasets

Let's break down why your original code struggles with large values of x (like the 8-digit DOB example) and fix it with practical, performance-focused changes.

Key Performance Bottlenecks in the Original Code

  1. Huge In-Memory Lists: Storing que and ans as lists with millions/billions of elements eats up massive RAM and slows down all subsequent operations. You don't actually need to track every question's options—you only care about how many each student got right/wrong.
  2. Triple Nested Loops: Iterating over every category, every student, and every question leads to an O(total_questions * num_students) time complexity. For 10 million questions and 1000 students, that's 10 billion operations—way too slow.
  3. Redundant Calculations: Tracking noOfIncorrect incrementally for each wrong answer is unnecessary; we can calculate the total penalty using a mathematical formula instead of looping through each mistake.
  4. Buggy Logic: Your code has a critical error in how it tracks question ranges (a = n+1 instead of a += n), leading to overlapping category data. It also overwrites totMarks for every student, so you only get the last student's score.

Optimized Solution Using NumPy (Fastest Approach)

NumPy's vectorized operations are orders of magnitude faster than pure Python loops for large datasets. Here's the revamped code:

import numpy as np

# Input handling
user_input = int(input("Enter a DOB:"))
ratio = [10, 4, 3, 2, 1]
total_ratio = sum(ratio)

# Calculate number of questions per category (matches original int conversion)
n_ques = [int((r / total_ratio) * user_input) for r in ratio]
num_students = 1000  # Corrected to match your problem statement (1000 candidates)

# Initialize storage for each category's scores
category_scores = [{} for _ in range(5)]

for cat_idx, num_questions in enumerate(n_ques):
    if num_questions == 0:
        continue  # Skip empty categories to avoid errors
    
    # Generate all students' results for this category in one go (vectorized)
    # 1 = correct, 0 = incorrect; shape: (num_students, num_questions)
    student_results = np.random.randint(0, 2, size=(num_students, num_questions))
    
    for stud_id in range(num_students):
        correct_count = np.sum(student_results[stud_id])
        incorrect_count = num_questions - correct_count
        
        # Calculate penalty using arithmetic series formula: 0.5*(1+2+...+k) = k*(k+1)/4
        penalty = incorrect_count * (incorrect_count + 1) / 4
        total_score = correct_count - penalty
        
        # Store score for this student in the category
        category_scores[cat_idx][stud_id + 1] = total_score

# Print results (adjust output as needed)
for i, scores in enumerate(category_scores, 1):
    print(f"Category {i} Scores: {scores}")

What Makes This Faster?

  • Vectorized Result Generation: NumPy generates all students' answers for a category in a single operation, avoiding millions of Python loop iterations.
  • No Unnecessary Storage: We don't save every question's options—only the count of correct/incorrect answers per student.
  • Mathematical Penalty Calculation: Instead of incrementing a counter for every wrong answer, we use the arithmetic series sum formula to compute total penalty in one step.
  • Fixed Range Tracking: No more overlapping category data, and each student's score is stored correctly.

Pure Python Alternative (No NumPy)

If you can't use NumPy, you can still optimize by leveraging statistical sampling and avoiding redundant loops:

import random

user_input = int(input("Enter a DOB:"))
ratio = [10, 4, 3, 2, 1]
total_ratio = sum(ratio)
n_ques = [int((r / total_ratio) * user_input) for r in ratio]
num_students = 1000

category_scores = [{} for _ in range(5)]

for cat_idx, num_questions in enumerate(n_ques):
    if num_questions == 0:
        continue
    
    for stud_id in range(num_students):
        # Count correct answers (each question has 1/4 chance of being correct)
        correct_count = sum(1 for _ in range(num_questions) if random.randint(1, 4) == random.randint(1, 4))
        incorrect_count = num_questions - correct_count
        
        # Same arithmetic series penalty calculation
        penalty = incorrect_count * (incorrect_count + 1) / 4
        total_score = correct_count - penalty
        
        category_scores[cat_idx][stud_id + 1] = total_score

for i, scores in enumerate(category_scores, 1):
    print(f"Category {i} Scores: {scores}")

This is slower than the NumPy version but still way faster than your original code, as it removes the huge que/ans lists and fixes the loop logic.

内容的提问来源于stack exchange,提问作者P N

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:56:42