You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

文件处理与字符串操作:对话文本分析及文件生成任务

任务处理方案

任务说明

现有对话脚本存储于文件conv.txt中,需完成以下两项技术任务:

  • 统计对话中唯一对话者的数量;
  • 为每个对话者创建对应的文本文件,将该角色所说的唯一单词逐行存入对应文件。

示例对话翻译

原文:
WILL: I’ve never seen wildlings do a thing like this. I’ve never seen a thing like this, not ever in my life.

WAYMAR ROYCE: How close did you get?

WILL: Close as any man would.

中文翻译:
威尔:我从没见过野人干出这种事。我这辈子都没见过这种事,从来没有。

韦玛·罗伊斯:你离得有多近?

威尔:近到任何人能达到的地步。

Python实现代码

import re
from collections import defaultdict

# 读取对话文件
with open('conv.txt', 'r', encoding='utf-8') as f:
    lines = [line.strip() for line in f if line.strip()]

# 解析对话内容,收集每个对话者的唯一单词
speaker_words = defaultdict(set)
for line in lines:
    match = re.match(r'^([A-Z\s]+):\s*(.*)$', line)
    if match:
        speaker = match.group(1).strip()
        content = match.group(2)
        # 提取单词,保留带撇号的形式(如I’ve),过滤无效字符
        words = re.findall(r'\b\w+\’?\w+\b', content)
        for word in words:
            speaker_words[speaker].add(word)

# 统计并输出唯一对话者数量
unique_speaker_num = len(speaker_words)
print(f"唯一对话者数量:{unique_speaker_num}")

# 为每个对话者生成对应文件
for speaker, unique_words in speaker_words.items():
    filename = f"{speaker.replace(' ', '_')}.txt"
    with open(filename, 'w', encoding='utf-8') as f:
        for word in sorted(unique_words):
            f.write(f"{word}\n")
    print(f"已生成文件:{filename}")

代码说明

  • 用正则表达式精准匹配对话行的"对话者: 内容"格式,避免解析错误;
  • 使用set存储单词,自动实现去重,确保每个单词仅保留一次;
  • 处理对话者名称中的空格,生成合法的文件名(如WAYMAR_ROYCE.txt);
  • 写入文件前对单词按字母排序,保持内容整洁。

内容的提问来源于stack exchange,提问作者Immortal Bhoot

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.08 00:55:14