You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提取电影脚本中指定角色的台词并实现统计功能

Got it, let's build out that character line retrieval tool you need, plus the word count functionality. Here's a complete, practical implementation tailored to your use case:

电影脚本角色台词提取与词数统计工具

核心功能实现

First, we'll start with a clear data structure for your extracted lines (using sample data to simulate your already-extracted script), then build the wordsspoken function, interactive logic, and word count feature:

完整代码示例

# 模拟你已经提取好的电影台词数据
# 每个元素是字典,存储角色名和对应台词
script = [
    {"character": "MIA", "line": "Wait, you're telling me we're on a movie set?"},
    {"character": "SEBASTIAN", "line": "Yep, welcome to Hollywood. Every corner's a scene here."},
    {"character": "MIA", "line": "This is way cooler than my barista shift."},
    {"character": "SEBASTIAN", "line": "Tell me about it. Now, wanna sneak in and watch them shoot?"},
    {"character": "MIA", "line": "Are you sure that's allowed?"}
]

def wordsspoken(script, character_name):
    # 忽略大小写匹配,避免输入误差
    target_char = character_name.strip().upper()
    # 过滤指定角色的所有台词
    character_lines = [entry["line"] for entry in script if entry["character"].strip().upper() == target_char]
    
    if not character_lines:
        print(f"Oops, no lines found for {character_name}. Double-check the spelling?")
        return []
    
    # 格式化输出台词
    print(f"\n=== {character_name}'s Full Lines ===")
    for num, line in enumerate(character_lines, 1):
        print(f"{num}. {line}")
    
    # 返回台词列表,方便后续拓展分析
    return character_lines

# 交互式查询逻辑,支持多次查询直到用户退出
while True:
    user_input = input("\nEnter a character name (type 'quit' to exit): ")
    if user_input.lower() == "quit":
        print("Thanks for using the tool! Bye.")
        break
    
    # 获取台词并自动统计词数
    lines = wordsspoken(script, user_input)
    if lines:
        total_words = sum(len(line.split()) for line in lines)
        print(f"\nTotal words spoken by {user_input}: {total_words}")

关键细节说明

  • 大小写容错:The function converts both the input name and script character names to uppercase, so users can input mia, Mia, or MIA and still get the correct lines.
  • 可拓展性:The function returns a list of lines, so you can build on it for more analysis—like counting high-frequency words, analyzing line length distribution, etc.—beyond just word count.
  • 友好提示:If the input character doesn't exist, it gives a clear message to avoid user confusion.

适配不同的脚本格式

If your extracted script is in plain text lines like CHARACTER: Line content, run this quick preprocessing to convert it to the unified dictionary structure:

# 预处理示例:从原始文本行转换为标准结构
raw_script = [
    "MIA: Wait, you're telling me we're on a movie set?",
    "SEBASTIAN: Yep, welcome to Hollywood. Every corner's a scene here.",
    # 更多原始台词行...
]

script = []
for line in raw_script:
    # 检查行是否符合角色+台词的格式
    if ": " in line:
        char_name, line_text = line.split(": ", 1)
        script.append({"character": char_name.strip(), "line": line_text.strip()})

内容的提问来源于stack exchange,提问作者tgtrmr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:22:55