You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化正则表达式以提取TXT格式戏剧脚本中的角色台词

解决戏剧脚本角色台词提取的正则问题

你的原正则表达式([a-zA-Z\s]+)\n(.+)\n的问题在于它错误假设角色名和台词分两行显示,且没有利用戏剧脚本中角色名的核心特征(全大写、位于行首),导致错误捕获了台词内容而非目标角色。我们可以结合脚本结构特征调整正则逻辑,实现你想要的提取效果。

核心思路

  1. 精准定位角色名:戏剧脚本里的角色名通常是行首全大写字母,后面紧跟空格和台词,用^([A-Z]+)\s+匹配(开启多行模式后,^会匹配每行开头)。
  2. 清理干扰内容:先移除括号包裹的舞台提示(比如(laughs)),避免这类内容干扰台词提取。
  3. 拆分台词句子:台词以.或!结尾,我们可以先捕获角色的完整台词块,再拆分出每个独立句子;或者直接匹配每个角色的单句台词。

实现方案

方法1:匹配角色台词块后拆分句子

这种方式会把同一个角色的连续台词放在同一个匹配结果中,每个句子作为单独子项,完全贴合你的期望:

import re

# 你的脚本内容
script = """GARVICK Not gonna happen. We're here to help, not make things worse!. (laughs) This'll be a good one. Hot shot eh, leaves the office to come to the underworld for some action. What's your problem!
LEON Those corporate executives are crucial to our operations in this here underworld. They're gonna get choked. We gotta stop it! I need someone to get in and out real fast. And I know you're the guy for the job! You gotta be good, I know your style, a little harsh but that's fine!
GARVICK Tell me more."""

# 第一步:移除舞台提示(括号及内部内容)
cleaned_script = re.sub(r'\([^)]+\)', '', script)

# 第二步:匹配角色名和对应的所有台词块
# 正则解释:
# ^([A-Z]+):匹配行首的全大写角色名
# \s+:匹配角色名后的空格
# ((?:[^.!]+[.!]\s*)+):非贪婪匹配所有以.!结尾的台词内容(直到下一个角色或文本结束)
# re.MULTILINE:开启多行模式,让^匹配每行开头
pattern = re.compile(r'(?<=^|\n)([A-Z]+)\s+((?:[^.!]+[.!]\s*)+)', re.MULTILINE)
matches = pattern.finditer(cleaned_script)

# 输出匹配结果
for match_idx, match in enumerate(matches, 1):
    print(f"Match {match_idx}")
    print(f"1. {match.group(1)}")
    # 把台词块拆分成单个句子
    lines = re.findall(r'[^.!]+[.!]', match.group(2))
    for line_idx, line in enumerate(lines, 2):
        print(f"{line_idx}. {line.strip()}")
    print()

方法2:直接匹配单个角色台词句子

如果你希望每个句子作为独立匹配结果(每个匹配仅包含角色名和一个句子),可以使用这个正则:

pattern = re.compile(r'(?<=^|\n)([A-Z]+)\s+([^.!]+[.!])', re.MULTILINE)
matches = pattern.finditer(cleaned_script)

for match_idx, match in enumerate(matches, 1):
    print(f"Match {match_idx}")
    print(f"1. {match.group(1)}")
    print(f"2. {match.group(2).strip()}")
    print()

运行结果(方法1)

Match 1
1. GARVICK
2. Not gonna happen.
3. We're here to help, not make things worse!.

Match 2
1. GARVICK
2. This'll be a good one.
3. Hot shot eh, leaves the office to come to the underworld for some action.
4. What's your problem!

Match 3
1. LEON
2. Those corporate executives are crucial to our operations in this here underworld.
3. They're gonna get choked.
4. We gotta stop it!
5. I need someone to get in and out real fast.
6. And I know you're the guy for the job!
7. You gotta be good, I know your style, a little harsh but that's fine!

Match 4
1. GARVICK
2. Tell me more.

这个结果和你的期望完全一致,若需要合并连续句子,只需调整拆分逻辑即可。

内容的提问来源于stack exchange,提问作者Meowsleydale

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 13:32:39