Windows下Python3.5实现歌德戏剧场次台词至嵌套列表的追加问题
Alright, let's break down how to build that nested list structure for Goethe's plays using Python 3.5 on Windows. Since the plays follow a consistent structure, we can automate parsing lines into (character, line) tuples and grouping them by scene. Here's a practical, step-by-step solution:
First, let's clarify the end goal: we want a top-level list where each element is a sublist representing one scene. Each sublist will contain tuples formatted as (角色名, 台词). For example:
[ [("FAUST", "Was ziehst du mich so hastig hin?"), ("MEPHISTOPHELES", "Ich will dich führen, wo du froh bist.")], # 场次1 [("GRETCHEN", "Mein lieber Herr, was macht ihr hier?"), ...] # 场次2 ]
Goethe's plays follow predictable formatting cues we can leverage:
- Scenes are usually marked with clear headers (like
SCENE 1orACT I, SCENE 2) - Character names are often in all caps, followed by a colon (e.g.,
FAUST:) - Lines belonging to a character follow immediately after their name
We'll use these cues to split the text into scenes and map lines to their respective characters.
Here's a code snippet tailored to your Windows environment and Python 3.5 constraints (no 3.6+ features like f-strings):
import os def parse_goethe_play(file_path): # Initialize our nested structure: top list = scenes, sublists = character-line tuples play_scenes = [] current_scene = [] current_character = None # Handle Windows file encoding (adjust to 'gbk' if your files use Chinese encoding) with open(file_path, 'r', encoding='utf-8') as play_file: for line in play_file: stripped_line = line.strip() # Skip empty lines if not stripped_line: continue # Detect scene changes (adjust this condition to match your play's scene headers) if stripped_line.startswith("SCENE") or stripped_line.startswith("ACT"): # Save the current scene if it has content, then reset for the next scene if current_scene: play_scenes.append(current_scene) current_scene = [] continue # Detect character names (assuming format: "CHARACTER:") if ':' in stripped_line and stripped_line.split(':')[0].isupper(): char_name = stripped_line.split(':')[0].strip() current_character = char_name # Check if the line has dialogue after the colon line_content = stripped_line.split(':', 1)[1].strip() if line_content: current_scene.append((current_character, line_content)) # If we have an active character, this line is their dialogue elif current_character: current_scene.append((current_character, stripped_line)) # Don't forget to add the last scene to the top-level list if current_scene: play_scenes.append(current_scene) return play_scenes # Example usage (Windows path uses raw string to avoid escape issues) faust_data = parse_goethe_play(r"C:\GoethePlays\Faust_Part1.txt") # Test: print the first 3 lines of the first scene print("First 3 lines of Scene 1:") for line_tuple in faust_data[0][:3]: print(f"{line_tuple[0]}: {line_tuple[1]}")
If your text files use slightly different formatting, adjust these parts:
- Scene detection: Modify the
startswithcondition to match your play's scene headers (e.g.,if "场次" in stripped_linefor Chinese-labeled scenes) - Character detection: If character names aren't all caps, use other cues like checking for a period instead of a colon, or that the line is short and doesn't look like dialogue
- Stage directions: Add a check to skip or tag stage directions (e.g.,
if stripped_line.startswith('['): continueto ignore bracketed directions)
- File paths: Use raw strings (like
r"C:\...") or forward slashes ("C:/...") to avoid escape character errors - Encoding: If you get a
UnicodeDecodeError, tryencoding='gbk'orencoding='utf-8-sig'(for files with a UTF-8 BOM) - Python 3.5 compatibility: Avoid using features introduced in 3.6+, like f-strings or the walrus operator (
:=)—stick to.format()for string formatting
内容的提问来源于stack exchange,提问作者FabianPeters

