You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python:移除重复列表项并删除原文档对应页面前三行

解决文档重复页面前三行的去重与删除问题

首先得帮你理清之前代码的问题:你用set(line_list[0]).intersection(line_list[2:])的思路完全偏了——列表是不可哈希的,没法直接放进set里当元素;而且intersection是找集合间的元素交集,不是找重复的整个子列表,这才导致输出空列表。

下面给你一套完整的解决方案,分步骤实现你的需求:

核心思路

  1. 提取每页前三行:先把文档按页面拆分,提取每页的前三行,保存为列表的列表(每个子列表对应一页的前三行)
  2. 识别重复的子列表:把每个子列表转成tuple(可哈希,能当字典键),用字典记录每个tuple的出现次数和对应页面索引,找出重复的页面
  3. 处理原文档:遍历原文档的每个页面,对重复页面删除其前三行,对唯一页面保留完整内容
  4. 写入新文档:把处理后的内容合并写入新文件

具体代码实现

情况1:文档按固定行数分页

如果你的文档每页是固定行数(比如每页10行),可以用这套代码:

def extract_three_lines_per_page(file_path, lines_per_page=10):
    """提取每页的前三行,同时返回原文档所有行和每页行数"""
    with open(file_path, 'r', encoding='utf-8') as f:
        original_lines = f.readlines()
        total_pages = (len(original_lines) + lines_per_page - 1) // lines_per_page
        pages_three_lines = []
        for page_idx in range(total_pages):
            start = page_idx * lines_per_page
            page_lines = original_lines[start:start+lines_per_page]
            # 取前三行,不足三行则取现有所有行
            three_lines = [line.strip() for line in page_lines[:3]]
            pages_three_lines.append(three_lines)
    return original_lines, pages_three_lines, lines_per_page

def find_duplicate_page_indices(pages_three_lines):
    """找出重复前三行对应的页面索引"""
    seen = {}
    duplicate_indices = set()
    for idx, line_list in enumerate(pages_three_lines):
        # 转成tuple才能作为字典键
        key = tuple(line_list)
        if key in seen:
            # 标记当前页面为重复页
            duplicate_indices.add(idx)
        else:
            seen[key] = idx
    return duplicate_indices

def process_and_save_document(original_lines, duplicate_indices, lines_per_page, output_path):
    """处理原文档,删除重复页的前三行,保存到新文件"""
    processed_lines = []
    total_pages = (len(original_lines) + lines_per_page - 1) // lines_per_page
    for page_idx in range(total_pages):
        start = page_idx * lines_per_page
        page_lines = original_lines[start:start+lines_per_page]
        if page_idx in duplicate_indices:
            # 删除前三行,保留剩余内容
            processed_lines.extend(page_lines[3:])
        else:
            # 保留完整页面
            processed_lines.extend(page_lines)
    # 写入新文件
    with open(output_path, 'w', encoding='utf-8') as f:
        f.writelines(processed_lines)

# 调用示例
if __name__ == "__main__":
    input_file = "your_original_doc.txt"
    output_file = "processed_doc.txt"
    original_lines, pages_three_lines, lines_per_page = extract_three_lines_per_page(input_file, lines_per_page=10)
    duplicate_indices = find_duplicate_page_indices(pages_three_lines)
    process_and_save_document(original_lines, duplicate_indices, lines_per_page, output_file)

情况2:文档按换页符(\f)分页

如果你的文档是用换页符分隔页面,修改拆分逻辑即可:

def split_document_by_page_break(file_path):
    """按换页符拆分文档,提取每页前三行"""
    with open(file_path, 'r', encoding='utf-8') as f:
        content = f.read()
        # 按换页符拆分页面
        original_pages = [page.splitlines(keepends=True) for page in content.split('\f')]
        # 提取每页前三行
        pages_three_lines = []
        for page in original_pages:
            three_lines = [line.strip() for line in page[:3]]
            pages_three_lines.append(three_lines)
    return original_pages, pages_three_lines

# 处理和保存逻辑调整为按页面处理
def process_pages(original_pages, duplicate_indices, output_path):
    processed_pages = []
    for idx, page in enumerate(original_pages):
        if idx in duplicate_indices:
            # 删除前三行
            processed_pages.append(page[3:])
        else:
            processed_pages.append(page)
    # 合并成完整内容写入文件
    with open(output_path, 'w', encoding='utf-8') as f:
        # 用换页符分隔页面
        f.write('\f'.join([''.join(page) for page in processed_pages]))

# 调用示例
if __name__ == "__main__":
    input_file = "your_original_doc.txt"
    output_file = "processed_doc.txt"
    original_pages, pages_three_lines = split_document_by_page_break(input_file)
    duplicate_indices = find_duplicate_page_indices(pages_three_lines)
    process_pages(original_pages, duplicate_indices, output_file)

关键细节说明

  • 为什么用tuple?因为列表是可变类型,不可哈希,没法作为字典的键;而tuple是不可变的,能正常被字典识别。
  • 去重逻辑:只保留第一次出现的子列表,后续重复的页面都会被标记为需要删除前三行的对象。
  • 兼容性:代码里处理了页面不足三行的情况,避免索引越界。

内容的提问来源于stack exchange,提问作者user11464178

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:43:18