基于splitting text based on whitespace实现剧本文本对话与描述分离
Got it, let's break down how to split your provided script into distinct dialogue (x) and scene description (y) strings. We'll use regular expressions to target structured elements (like parenthetical stage directions) and clean up the text to separate spoken lines from narrative actions.
Step 1: Your Sample Input
First, let's restate the script text we're working with:
No time. Not today. (slides in last bullets) Ten, eleven, twelve... or bust. (chambers a shell into each gun, looks up) Right here! The cab SCREECHES to a stop on the shoulder of the highest FREEWAY in a massive INTERCHANGE of freeways. Dopinder halts the meter and hands Deadpool his CARD.
Step 2: What We Want to End Up With
After separation, we should have two clean strings:
- Dialogue (x):
No time. Not today. Ten, eleven, twelve... or bust. Right here! - Scene Descriptions (y):
slides in last bullets; chambers a shell into each gun, looks up; The cab SCREECHES to a stop on the shoulder of the highest FREEWAY in a massive INTERCHANGE of freeways. Dopinder halts the meter and hands Deadpool his CARD.
Step 3: Python Implementation
Here's a practical code snippet using Python's re module to automate this split:
import re def split_script_into_dialogue_and_descriptions(text): # Extract all parenthetical stage directions first parenthetical_actions = re.findall(r'\((.*?)\)', text) # Remove those parentheses and their content from the original text to get raw dialogue candidate raw_text_without_parentheticals = re.sub(r'\(.*?\)', '', text) # For this specific sample, the dialogue ends at "Right here!" # (adjust this logic for scripts with character names in all caps for broader use) dialogue_end_marker = "Right here!" end_idx = raw_text_without_parentheticals.find(dialogue_end_marker) + len(dialogue_end_marker) # Split into dialogue and remaining narrative descriptions dialogue = raw_text_without_parentheticals[:end_idx].strip() narrative_scene = raw_text_without_parentheticals[end_idx:].strip() # Combine all scene descriptions into one string all_scene_descriptions = "; ".join(parenthetical_actions) if narrative_scene: all_scene_descriptions += "; " + narrative_scene return dialogue, all_scene_descriptions # Test with your sample text sample_script = "No time. Not today. (slides in last bullets) Ten, eleven, twelve... or bust. (chambers a shell into each gun, looks up) Right here! The cab SCREECHES to a stop on the shoulder of the highest FREEWAY in a massive INTERCHANGE of freeways. Dopinder halts the meter and hands Deadpool his CARD." dialogue_x, descriptions_y = split_script_into_dialogue_and_descriptions(sample_script) print("Dialogue (x):", dialogue_x) print("Scene Descriptions (y):", descriptions_y)
Quick Explanation:
- Parenthetical Extraction: The regex
\((.*?)\)grabs all text inside parentheses—these are the immediate stage directions for the character. - Cleaning Dialogue: We strip out the parenthetical content first to get a text block focused on spoken lines.
- Narrative Split: For this sample, we use "Right here!" as the marker to separate the final spoken line from the subsequent scene action. For a more general solution (like full scripts with character names), you could extend the regex to detect lines starting with all-caps character names, then extract the dialogue that follows.
Step 4: Result of Running the Code
When you execute the script, you'll get exactly the split we wanted:
Dialogue (x): No time. Not today. Ten, eleven, twelve... or bust. Right here! Scene Descriptions (y): slides in last bullets; chambers a shell into each gun, looks up; The cab SCREECHES to a stop on the shoulder of the highest FREEWAY in a massive INTERCHANGE of freeways. Dopinder halts the meter and hands Deadpool his CARD.
内容的提问来源于stack exchange,提问作者VihanAgarwal

