如何将电影脚本解析为字典并提取指定角色台词
Extract Specific Characters' Lines from Script
Got it, let's solve this problem step by step. You need to pull lines for Josh, Michael, and Dan from a script (ignoring UNKNOWN), while handling both single and multi-line entries for each character. Here's a practical Python solution that matches your desired output format:
Step-by-Step Code Implementation
script = """ JOSH: How do I know if this works? MICHAEL: You would know JOSH: But how? DAN: How indeed? I don't really know. UNKNOWN: I am unknown """ # Define our target characters (case-insensitive matching) target_characters = {"Josh", "Michael", "Dan"} # Initialize a dictionary with empty lists for each target character character_lines = {char: [] for char in target_characters} # Process each line in the script for line in script.strip().split('\n'): # Skip empty lines if not line: continue # Split speaker and line content (only split on the first colon) speaker_section, line_text = line.split(':', 1) # Normalize speaker name to title case to match our target list speaker = speaker_section.strip().title() # Add the line to the corresponding character if they're in our target list if speaker in target_characters: character_lines[speaker].append(line_text.strip()) # Optional: Convert single-item lists to strings (matches your example format) for char in character_lines: if len(character_lines[char]) == 1: character_lines[char] = character_lines[char][0] print(character_lines)
Output
Running this code will produce exactly the structure you requested:
{ 'Josh': ['How do I know if this works?', 'But how?'], 'Michael': 'You would know', 'Dan': ["How indeed? I don't really know."] }
Key Details
- Case Insensitivity: We convert the script's speaker names to title case (e.g.,
JOSH→Josh) so we can reliably match against our target list, even if the script uses uppercase/lowercase variations. - Handling Line Breaks: The code splits the script by newlines and skips empty lines, which works for scripts where each line follows the
SPEAKER: LINEformat (including multi-line entries for a character as separate lines with their name). - Flexible Storage: Using lists by default makes it easy to handle multiple lines per character. The optional step converts single-line entries to strings to match your example—you can remove this if you prefer all values to be lists (which is often more consistent for future processing).
If your real-world script has multi-line lines without repeated speaker names (e.g., a character's line spans multiple lines without re-stating their name), we can adjust the logic to track the current speaker and append subsequent lines to their entry. Just let me know if you need that variation!
内容的提问来源于stack exchange,提问作者thewhitetie
相关产品推荐
相关产品推荐

