如何扩展Son Huang的Python文本分割方案实现内容迁移与时间戳添加?
问题描述
我有一段需要分割为多文本块并写入不同文件的文本,打算基于l'mahdi的方案,采用Son Huang的实现思路。现在输入文本已做修改:所有以note::开头的行在逗号前新增了内容,且每个文本块都新增了highlight::开头的行。输入文本如下:
company:: acme products department:: sales floor:: 1 name:: Joe Blogs phone:: 123456789 email:: joeblogs@email.com address:: 123 Main Street note:: highlight text, blah blah blah timestamp:: highlight:: name:: Josephine Blogs phone:: 43217890 email:: josephineblogs@email.com address:: 123 Main Street note:: Another highlight here, More blah blah timestamp:: highlight:: name:: John Smith phone:: 23498689 email:: johnsmith@email.com address:: 1 North Street note:: Amazing text, Some more blah timestamp:: highlight::
需要在Son Huang的解决方案基础上添加哪些内容,才能得到如下期望输出?可以看到note::行逗号前的文本移到了highlight::行(逗号已移除),同时为timestamp::行添加递增时间戳,每个文本块生成独立文件并在末尾附加顶部的公司信息:
# chunk_1.txt name:: Joe Blogs phone:: 123456789 email:: joeblogs@email.com address:: 123 Main Street note:: blah blah blah timestamp:: 2022-08-07 (13h 10m 08s) highlight:: highlight text company:: acme products department:: sales floor:: 1 # chunk_2.txt name:: Josephine Blogs phone:: 43217890 email:: josephineblogs@email.com address:: 123 Main Street note:: More blah blah timestamp:: 2022-08-07 (13h 10m 09s) highlight:: Another highlight here company:: acme products department:: sales floor:: 1 # chunk_3.txt name:: John Smith phone:: 23498689 email:: johnsmith@email.com address:: 1 North Street note:: Some more blah timestamp:: 2022-08-07 (13h 10m 10s) highlight:: Amazing text company:: acme products department:: sales floor:: 1
所需添加的内容与实现步骤
要实现需求,需要在原有方案基础上添加/修改以下核心逻辑:
1. 提取并保存顶部公共公司信息
读取输入文本时,先提取开头的company::、department::、floor::行作为公共内容,后续每个文件都要附加这部分内容。
2. 分割文本块
以空行为分隔符,将输入文本拆分为多个员工信息块,跳过开头的公共信息块。
3. 处理每个文本块的核心操作
对每个员工信息块,完成三个关键处理:
- 拆分
note::行:将note::行中逗号前的内容移至highlight::行,逗号后的内容保留为新的note::行文本。 - 生成递增时间戳:基于起始时间,每个块的时间戳依次递增1秒,格式化为
YYYY-MM-DD (HHh MMm SSs)。 - 整合内容:按期望顺序组合处理后的员工信息与公共公司信息。
4. 写入独立文件
将每个处理完成的文本块写入对应的chunk_N.txt文件。
完整代码示例
import datetime # 输入文本(可替换为从文件读取) input_text = """company:: acme products department:: sales floor:: 1 name:: Joe Blogs phone:: 123456789 email:: joeblogs@email.com address:: 123 Main Street note:: highlight text, blah blah blah timestamp:: highlight:: name:: Josephine Blogs phone:: 43217890 email:: josephineblogs@email.com address:: 123 Main Street note:: Another highlight here, More blah blah timestamp:: highlight:: name:: John Smith phone:: 23498689 email:: johnsmith@email.com address:: 1 North Street note:: Amazing text, Some more blah timestamp:: highlight::""" # 1. 提取公共公司信息 lines = input_text.split('\n') common_info = [] for i, line in enumerate(lines): if line.strip() == '': break common_info.append(line.strip()) # 跳过空行,获取员工块起始位置 employee_blocks_start = i + 1 # 2. 分割员工信息块 employee_blocks = [] current_block = [] for line in lines[employee_blocks_start:]: if line.strip() == '': if current_block: employee_blocks.append(current_block) current_block = [] else: current_block.append(line.strip()) # 添加最后一个块 if current_block: employee_blocks.append(current_block) # 3. 设置起始时间戳,每个块递增1秒 start_time = datetime.datetime(2022, 8, 7, 13, 10, 8) # 4. 处理每个块并写入文件 for idx, block in enumerate(employee_blocks, 1): processed_lines = [] highlight_content = "" note_content = "" timestamp_line = "" for line in block: if line.startswith('note::'): # 拆分note内容 _, note_full = line.split('::', 1) if ',' in note_full: highlight_part, note_part = note_full.split(',', 1) highlight_content = highlight_part.strip() note_content = note_part.strip() processed_lines.append(f'note:: {note_content}') else: # 处理无逗号的情况(按需求默认保留原内容) processed_lines.append(line) highlight_content = "" elif line.startswith('timestamp::'): # 生成递增时间戳 current_time = start_time + datetime.timedelta(seconds=idx-1) timestamp_str = current_time.strftime("%Y-%m-%d (%Hh %Mm %Ss)") timestamp_line = f'timestamp:: {timestamp_str}' elif line.startswith('highlight::'): # 替换highlight内容 processed_lines.append(f'highlight:: {highlight_content}') else: # 保留其他行 processed_lines.append(line) # 添加时间戳行 processed_lines.append(timestamp_line) # 添加公共信息到末尾 processed_lines.extend(common_info) # 生成文件内容 file_content = '\n'.join(processed_lines) # 写入文件 filename = f'chunk_{idx}.txt' with open(filename, 'w') as f: f.write(file_content)
内容的提问来源于stack exchange,提问作者supposer
相关产品推荐
相关产品推荐

