You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何统计文本文件各段落行数并移除行数不符的末段

Simplify Paragraph Line Count & Removal Task

Got it, let's break this down into a clean, concise solution. First, let's align on the core requirements (based on your example and description):

  • Split your text file into paragraphs (standardly separated by one or more blank lines)
  • Count how many lines are in each paragraph
  • Remove the last paragraph if its line count doesn't match the pattern of the other paragraphs (like your example where the last segment is shorter)
  • Output or write back the remaining content

Your current approach tracks line counts but doesn't handle paragraph separation, which is why it feels complicated. Here's a far more straightforward Python implementation:

import re

def clean_paragraphs(file_path):
    # Read the entire file in one go (simple for most text files)
    with open(file_path, 'r') as f:
        full_content = f.read()
    
    # Split content into paragraphs, ignoring empty/whitespace-only segments
    # Regex handles multiple blank lines between paragraphs
    paragraphs = [para.strip() for para in re.split(r'\n\s*\n', full_content) if para.strip()]
    
    # Edge case: if there's only one paragraph, no need to process
    if len(paragraphs) <= 1:
        return full_content
    
    # Count lines in each paragraph
    line_counts = [len(para.splitlines()) for para in paragraphs]
    
    # Define your condition: remove last paragraph if its line count differs from the first
    # Adjust this condition if you need a different logic (e.g., compare to majority)
    if line_counts[-1] != line_counts[0]:
        processed_paragraphs = paragraphs[:-1]
    else:
        processed_paragraphs = paragraphs
    
    # Join paragraphs back with standard blank line separators
    return '\n\n'.join(processed_paragraphs)

# Example usage:
# result = clean_paragraphs("your_text_file.txt")
# with open("processed_file.txt", 'w') as f:
#     f.write(result)

Why this works better:

  • Robust paragraph splitting: The regex r'\n\s*\n' handles any number of blank lines between paragraphs (not just one), which is common in real-world text files.
  • Readable logic: Every step is explicit—split, count, judge, recombine—no confusing index tracking.
  • Flexible condition: If you need to change the rule (e.g., remove last paragraph if it's shorter than average), just tweak the if line_counts[-1] != line_counts[0] line to fit your needs.

For your example input (assuming paragraphs are separated by a blank line):

Original file content:

black yellow pink
hills mountain liver

barbecue spaghetti

The first paragraph has 2 lines, the last has 1. The code will remove the last paragraph, giving you the expected output:

black yellow pink
hills mountain liver

内容的提问来源于stack exchange,提问作者meme

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 12:20:27