You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何扩展Son Huang的Python文本分割方案实现内容迁移与时间戳添加?

问题描述

我有一段需要分割为多文本块并写入不同文件的文本,打算基于l'mahdi的方案,采用Son Huang的实现思路。现在输入文本已做修改:所有以note::开头的行在逗号前新增了内容,且每个文本块都新增了highlight::开头的行。输入文本如下:

company:: acme products
department:: sales
floor:: 1

name:: Joe Blogs 
phone:: 123456789
email:: joeblogs@email.com
address:: 123 Main Street
note:: highlight text, blah blah blah
timestamp::
highlight::

name:: Josephine Blogs 
phone:: 43217890
email:: josephineblogs@email.com
address:: 123 Main Street
note:: Another highlight here, More blah blah
timestamp::
highlight::

name:: John Smith 
phone:: 23498689
email:: johnsmith@email.com
address:: 1 North Street
note:: Amazing text, Some more blah
timestamp::
highlight::

需要在Son Huang的解决方案基础上添加哪些内容,才能得到如下期望输出?可以看到note::行逗号前的文本移到了highlight::行(逗号已移除),同时为timestamp::行添加递增时间戳,每个文本块生成独立文件并在末尾附加顶部的公司信息:

# chunk_1.txt

name:: Joe Blogs
phone:: 123456789
email:: joeblogs@email.com
address:: 123 Main Street
note:: blah blah blah
timestamp:: 2022-08-07 (13h 10m 08s)
highlight:: highlight text
company:: acme products
department:: sales
floor:: 1

# chunk_2.txt

name:: Josephine Blogs
phone:: 43217890
email:: josephineblogs@email.com
address:: 123 Main Street
note:: More blah blah
timestamp:: 2022-08-07 (13h 10m 09s)
highlight:: Another highlight here
company:: acme products
department:: sales
floor:: 1

# chunk_3.txt

name:: John Smith
phone:: 23498689
email:: johnsmith@email.com
address:: 1 North Street
note:: Some more blah
timestamp:: 2022-08-07 (13h 10m 10s)
highlight:: Amazing text
company:: acme products
department:: sales
floor:: 1
所需添加的内容与实现步骤

要实现需求,需要在原有方案基础上添加/修改以下核心逻辑:

1. 提取并保存顶部公共公司信息

读取输入文本时,先提取开头的company::、department::、floor::行作为公共内容,后续每个文件都要附加这部分内容。

2. 分割文本块

以空行为分隔符,将输入文本拆分为多个员工信息块,跳过开头的公共信息块。

3. 处理每个文本块的核心操作

对每个员工信息块,完成三个关键处理:

  • 拆分note::行:将note::行中逗号前的内容移至highlight::行,逗号后的内容保留为新的note::行文本。
  • 生成递增时间戳:基于起始时间,每个块的时间戳依次递增1秒,格式化为YYYY-MM-DD (HHh MMm SSs)。
  • 整合内容:按期望顺序组合处理后的员工信息与公共公司信息。

4. 写入独立文件

将每个处理完成的文本块写入对应的chunk_N.txt文件。

完整代码示例

import datetime

# 输入文本(可替换为从文件读取)
input_text = """company:: acme products
department:: sales
floor:: 1

name:: Joe Blogs 
phone:: 123456789
email:: joeblogs@email.com
address:: 123 Main Street
note:: highlight text, blah blah blah
timestamp::
highlight::

name:: Josephine Blogs 
phone:: 43217890
email:: josephineblogs@email.com
address:: 123 Main Street
note:: Another highlight here, More blah blah
timestamp::
highlight::

name:: John Smith 
phone:: 23498689
email:: johnsmith@email.com
address:: 1 North Street
note:: Amazing text, Some more blah
timestamp::
highlight::"""

# 1. 提取公共公司信息
lines = input_text.split('\n')
common_info = []
for i, line in enumerate(lines):
    if line.strip() == '':
        break
    common_info.append(line.strip())
# 跳过空行,获取员工块起始位置
employee_blocks_start = i + 1

# 2. 分割员工信息块
employee_blocks = []
current_block = []
for line in lines[employee_blocks_start:]:
    if line.strip() == '':
        if current_block:
            employee_blocks.append(current_block)
            current_block = []
    else:
        current_block.append(line.strip())
# 添加最后一个块
if current_block:
    employee_blocks.append(current_block)

# 3. 设置起始时间戳,每个块递增1秒
start_time = datetime.datetime(2022, 8, 7, 13, 10, 8)

# 4. 处理每个块并写入文件
for idx, block in enumerate(employee_blocks, 1):
    processed_lines = []
    highlight_content = ""
    note_content = ""
    timestamp_line = ""
    for line in block:
        if line.startswith('note::'):
            # 拆分note内容
            _, note_full = line.split('::', 1)
            if ',' in note_full:
                highlight_part, note_part = note_full.split(',', 1)
                highlight_content = highlight_part.strip()
                note_content = note_part.strip()
                processed_lines.append(f'note:: {note_content}')
            else:
                # 处理无逗号的情况(按需求默认保留原内容)
                processed_lines.append(line)
                highlight_content = ""
        elif line.startswith('timestamp::'):
            # 生成递增时间戳
            current_time = start_time + datetime.timedelta(seconds=idx-1)
            timestamp_str = current_time.strftime("%Y-%m-%d (%Hh %Mm %Ss)")
            timestamp_line = f'timestamp:: {timestamp_str}'
        elif line.startswith('highlight::'):
            # 替换highlight内容
            processed_lines.append(f'highlight:: {highlight_content}')
        else:
            # 保留其他行
            processed_lines.append(line)
    # 添加时间戳行
    processed_lines.append(timestamp_line)
    # 添加公共信息到末尾
    processed_lines.extend(common_info)
    
    # 生成文件内容
    file_content = '\n'.join(processed_lines)
    # 写入文件
    filename = f'chunk_{idx}.txt'
    with open(filename, 'w') as f:
        f.write(file_content)

内容的提问来源于stack exchange,提问作者supposer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 14:30:25