You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何挖掘与学生高低成绩相关的行为文本特征词汇?

Hey Sara, nice work knocking out the preprocessing and TF-IDF setup—you’re already in a great spot to uncover those grade-linked features! Let’s walk through practical, actionable steps to connect your ngrams to high and low scores:

1. Define Clear Grade Labels First

Before diving into feature analysis, you need to turn raw grades into a binary (or categorical) label set so you can compare groups. For example:

  • Split grades into high score (top 20-30% of your dataset) and low score (bottom 20-30%)—this avoids noise from middle-of-the-road scores.
  • Or use a pass/fail threshold if that makes more sense for your data (e.g., grades ≥80 = high, <60 = low).
  • Store these labels in a variable like y (e.g., y = 1 for high scores, y = 0 for low scores).

2. Analyze Feature-Grade Relationships

Here are 3 robust methods to find which ngrams correlate with high/low scores:

Method 1: Frequency Ratio (Quick & Intuitive)

This is the simplest way to spot trends—calculate how much more often an ngram appears in one group vs. the other:

  1. For each ngram, count its total occurrences in the high-score group (count_high) and low-score group (count_low).
  2. Compute the ratio: (count_high / total_high_docs) / (count_low / total_low_docs)
    • A ratio > 2 means the ngram is far more common in high-score texts.
    • A ratio < 0.5 means it’s far more common in low-score texts.
  3. Example code snippet (using pandas):
    import pandas as pd
    
    # Assuming you have a DataFrame with 'text', 'grade', and 'label' columns
    high_docs = df[df['label'] == 1]['text']
    low_docs = df[df['label'] == 0]['text']
    
    # Count ngram frequencies (reuse your TF-IDF vocabulary here)
    from sklearn.feature_extraction.text import CountVectorizer
    vectorizer = CountVectorizer(vocabulary=your_tfidf_vocab, ngram_range=(1,2))
    high_counts = vectorizer.fit_transform(high_docs).sum(axis=0).A1
    low_counts = vectorizer.fit_transform(low_docs).sum(axis=0).A1
    
    # Calculate ratios
    freq_ratios = (high_counts / len(high_docs)) / (low_counts / len(low_docs))
    ngram_ratios = pd.Series(freq_ratios, index=vectorizer.get_feature_names_out())
    
    # Get top high/low linked ngrams
    top_high_ngrams = ngram_ratios.sort_values(ascending=False).head(10)
    top_low_ngrams = ngram_ratios.sort_values(ascending=True).head(10)
    

Method 2: Mutual Information (Statistically Rigorous)

Mutual Information measures how much knowing an ngram reduces uncertainty about a student’s grade—it’s perfect for sparse text data like TF-IDF. Use scikit-learn’s implementation:

from sklearn.feature_selection import mutual_info_classif

# X_tfidf is your precomputed TF-IDF feature matrix
mi_scores = mutual_info_classif(X_tfidf, y, random_state=42)

# Map scores to ngrams
mi_series = pd.Series(mi_scores, index=your_tfidf_vocab)

# Sort to find top correlated features
top_high_mi = mi_series.sort_values(ascending=False).head(10)
top_low_mi = mi_series[mi_series > 0].sort_values(ascending=True).head(10)

Higher mutual information scores mean the ngram is more predictive of grade group.

Method 3: Model Feature Importance (Predictive Validation)

Train a simple classification model (like Logistic Regression) and use its coefficients to see which ngrams drive predictions for high/low scores:

from sklearn.linear_model import LogisticRegression

lr = LogisticRegression(max_iter=1000)
lr.fit(X_tfidf, y)

# Extract coefficients (positive = linked to high scores, negative = linked to low)
feature_coeffs = pd.Series(lr.coef_[0], index=your_tfidf_vocab)

top_high_coeffs = feature_coeffs.sort_values(ascending=False).head(10)
top_low_coeffs = feature_coeffs.sort_values(ascending=True).head(10)

This method confirms that your features aren’t just statistically correlated—they’re actually useful for predicting grades.

3. Visualize Your Findings

Make it easy to interpret by plotting top features:

import seaborn as sns
import matplotlib.pyplot as plt

# Plot top high-score ngrams
plt.figure(figsize=(10,6))
sns.barplot(x=top_high_coeffs.values, y=top_high_coeffs.index)
plt.title("Top Ngrams Linked to High Student Scores")
plt.xlabel("Logistic Regression Coefficient")
plt.show()

# Plot top low-score ngrams
plt.figure(figsize=(10,6))
sns.barplot(x=top_low_coeffs.values, y=top_low_coeffs.index)
plt.title("Top Ngrams Linked to Low Student Scores")
plt.xlabel("Logistic Regression Coefficient")
plt.show()

4. Clean Up Noise

  • Filter out ngrams that appear in fewer than 5-10 documents—rare terms can skew results and aren’t generalizable.
  • Check for redundant ngrams (e.g., if "study" and "studying" both appear, your stemming should have handled this, but double-check).

By combining these methods—starting with quick frequency checks, validating with statistical tests, and confirming with model importance—you’ll get a solid list of grade-linked features like you expected (e.g., "study", "completed homework" for high scores; "missed class", "late assignment" for low scores).

内容的提问来源于stack exchange,提问作者sara

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:43:54