基于Python的小说文本视角分析技术需求:统计角色视角词数
Got it, let's break down how to implement a Point of View (PoV) analysis for novels using Python, focused on counting word counts per character's perspective—just like the approach in Statistical Analysis of WoT. Here's a step-by-step, practical method:
First, you need your novel in a machine-readable format (plain text .txt is ideal). If you're working with a published novel like The Wheel of Time, you might find pre-processed versions with PoV markers online; for your own work or less common books, you'll need to prepare the text yourself.
- Clean the text: Remove any headers, footers, or formatting artifacts (like page numbers) that aren't part of the narrative.
- Save as UTF-8: Ensure the file uses a universal encoding to avoid character issues.
To count words per character, you first need to identify which parts of the text are from each character's PoV. There are two main approaches:
Manual Labeling (Most Accurate)
Add explicit markers in your text to denote PoV shifts. For example:
[POV: Rand al'Thor]
The wind howled through the streets of Emond's Field, and Rand gripped his father's axe tighter...[POV: Egwene al'Vere]
Egwene watched Rand from the window, her fingers hovering over her weaving...
This is the method used in Statistical Analysis of WoT because it eliminates ambiguity. For long novels, you can use a text editor with search/replace to speed up labeling if PoV shifts are tied to chapter openings.
Automatic PoV Detection (Advanced)
If manual labeling is too time-consuming, you can use NLP models to predict PoV:
- Use Named Entity Recognition (NER) to identify character names in each paragraph.
- Train a classifier (e.g., fine-tuned BERT) on a small labeled subset of your novel to predict PoV for unlabeled sections.
- Note: This will have some error margin, so manual validation is recommended.
Once your text is labeled, use Python to extract each PoV section and count the words. Here's a straightforward script:
import re from collections import defaultdict def count_pov_word_counts(text_path): # Initialize a dictionary to hold word counts per character pov_word_counts = defaultdict(int) current_pov = None # Regex pattern to match our PoV markers (adjust if you used a different format) pov_pattern = re.compile(r'\[POV: (.*?)\]') with open(text_path, 'r', encoding='utf-8') as f: for line in f: # Check if the line contains a PoV marker pov_match = pov_pattern.search(line) if pov_match: current_pov = pov_match.group(1).strip() continue # If we're in a PoV section, count the words if current_pov: # Split line into words (simple split; adjust for better tokenization if needed) words = line.strip().split() # Exclude empty lines if words: pov_word_counts[current_pov] += len(words) return dict(pov_word_counts) # Example usage if __name__ == "__main__": word_counts = count_pov_word_counts("wheel_of_time_book1.txt") # Print results sorted by word count (descending) for char, count in sorted(word_counts.items(), key=lambda x: x[1], reverse=True): print(f"{char}: {count} words")
Improvements to Word Counting
- Better tokenization: Use
nltk.word_tokenize()instead ofsplit()to handle punctuation correctly (e.g., "don't" becomes one word instead of two). Install NLTK first withpip install nltk, then download the tokenizer:import nltk nltk.download('punkt') from nltk.tokenize import word_tokenize # Replace the word counting line with: words = word_tokenize(line.strip()) pov_word_counts[current_pov] += len(words) - Exclude stopwords: If you want to count meaningful words only, filter out common stopwords (e.g., "the", "and") using NLTK's stopword list:
from nltk.corpus import stopwords nltk.download('stopwords') stop_words = set(stopwords.words('english')) # Adjust the word counting: words = [word for word in word_tokenize(line.strip()) if word.lower() not in stop_words] pov_word_counts[current_pov] += len(words)
Once you have the word counts, you can extend the analysis to match the depth of Statistical Analysis of WoT:
- Calculate PoV percentage: For each character, compute what percentage of the total novel is from their perspective.
- Visualize results: Use
matplotlibto create a bar chart showing word counts per character:import matplotlib.pyplot as plt chars = list(word_counts.keys()) counts = list(word_counts.values()) plt.figure(figsize=(10, 6)) plt.bar(chars, counts, color='skyblue') plt.xlabel('PoV Characters') plt.ylabel('Total Words') plt.title('Novel PoV Word Count Distribution') plt.xticks(rotation=45, ha='right') plt.tight_layout() plt.show() - Lexical diversity: Calculate the number of unique words per character to compare their narrative style.
- For very large novels, consider processing the text in chunks to avoid memory issues.
- If using automatic PoV detection, start with a small labeled dataset to train your model—this will improve accuracy.
- The manual labeling method is the gold standard for academic analysis (like the WoT study) because it ensures no misclassification of PoV sections.
内容的提问来源于stack exchange,提问作者Vin Venture

