如何提取电影脚本中指定角色的台词并实现统计功能
Got it, let's build out that character line retrieval tool you need, plus the word count functionality. Here's a complete, practical implementation tailored to your use case:
电影脚本角色台词提取与词数统计工具
核心功能实现
First, we'll start with a clear data structure for your extracted lines (using sample data to simulate your already-extracted script), then build the wordsspoken function, interactive logic, and word count feature:
完整代码示例
# 模拟你已经提取好的电影台词数据 # 每个元素是字典,存储角色名和对应台词 script = [ {"character": "MIA", "line": "Wait, you're telling me we're on a movie set?"}, {"character": "SEBASTIAN", "line": "Yep, welcome to Hollywood. Every corner's a scene here."}, {"character": "MIA", "line": "This is way cooler than my barista shift."}, {"character": "SEBASTIAN", "line": "Tell me about it. Now, wanna sneak in and watch them shoot?"}, {"character": "MIA", "line": "Are you sure that's allowed?"} ] def wordsspoken(script, character_name): # 忽略大小写匹配,避免输入误差 target_char = character_name.strip().upper() # 过滤指定角色的所有台词 character_lines = [entry["line"] for entry in script if entry["character"].strip().upper() == target_char] if not character_lines: print(f"Oops, no lines found for {character_name}. Double-check the spelling?") return [] # 格式化输出台词 print(f"\n=== {character_name}'s Full Lines ===") for num, line in enumerate(character_lines, 1): print(f"{num}. {line}") # 返回台词列表,方便后续拓展分析 return character_lines # 交互式查询逻辑,支持多次查询直到用户退出 while True: user_input = input("\nEnter a character name (type 'quit' to exit): ") if user_input.lower() == "quit": print("Thanks for using the tool! Bye.") break # 获取台词并自动统计词数 lines = wordsspoken(script, user_input) if lines: total_words = sum(len(line.split()) for line in lines) print(f"\nTotal words spoken by {user_input}: {total_words}")
关键细节说明
- 大小写容错:The function converts both the input name and script character names to uppercase, so users can input
mia,Mia, orMIAand still get the correct lines. - 可拓展性:The function returns a list of lines, so you can build on it for more analysis—like counting high-frequency words, analyzing line length distribution, etc.—beyond just word count.
- 友好提示:If the input character doesn't exist, it gives a clear message to avoid user confusion.
适配不同的脚本格式
If your extracted script is in plain text lines like CHARACTER: Line content, run this quick preprocessing to convert it to the unified dictionary structure:
# 预处理示例:从原始文本行转换为标准结构 raw_script = [ "MIA: Wait, you're telling me we're on a movie set?", "SEBASTIAN: Yep, welcome to Hollywood. Every corner's a scene here.", # 更多原始台词行... ] script = [] for line in raw_script: # 检查行是否符合角色+台词的格式 if ": " in line: char_name, line_text = line.split(": ", 1) script.append({"character": char_name.strip(), "line": line_text.strip()})
内容的提问来源于stack exchange,提问作者tgtrmr
相关产品推荐
相关产品推荐

