如何统计文本文件各段落行数并移除行数不符的末段
Simplify Paragraph Line Count & Removal Task
Got it, let's break this down into a clean, concise solution. First, let's align on the core requirements (based on your example and description):
- Split your text file into paragraphs (standardly separated by one or more blank lines)
- Count how many lines are in each paragraph
- Remove the last paragraph if its line count doesn't match the pattern of the other paragraphs (like your example where the last segment is shorter)
- Output or write back the remaining content
Your current approach tracks line counts but doesn't handle paragraph separation, which is why it feels complicated. Here's a far more straightforward Python implementation:
import re def clean_paragraphs(file_path): # Read the entire file in one go (simple for most text files) with open(file_path, 'r') as f: full_content = f.read() # Split content into paragraphs, ignoring empty/whitespace-only segments # Regex handles multiple blank lines between paragraphs paragraphs = [para.strip() for para in re.split(r'\n\s*\n', full_content) if para.strip()] # Edge case: if there's only one paragraph, no need to process if len(paragraphs) <= 1: return full_content # Count lines in each paragraph line_counts = [len(para.splitlines()) for para in paragraphs] # Define your condition: remove last paragraph if its line count differs from the first # Adjust this condition if you need a different logic (e.g., compare to majority) if line_counts[-1] != line_counts[0]: processed_paragraphs = paragraphs[:-1] else: processed_paragraphs = paragraphs # Join paragraphs back with standard blank line separators return '\n\n'.join(processed_paragraphs) # Example usage: # result = clean_paragraphs("your_text_file.txt") # with open("processed_file.txt", 'w') as f: # f.write(result)
Why this works better:
- Robust paragraph splitting: The regex
r'\n\s*\n'handles any number of blank lines between paragraphs (not just one), which is common in real-world text files. - Readable logic: Every step is explicit—split, count, judge, recombine—no confusing index tracking.
- Flexible condition: If you need to change the rule (e.g., remove last paragraph if it's shorter than average), just tweak the
if line_counts[-1] != line_counts[0]line to fit your needs.
For your example input (assuming paragraphs are separated by a blank line):
Original file content:
black yellow pink hills mountain liver barbecue spaghettiThe first paragraph has 2 lines, the last has 1. The code will remove the last paragraph, giving you the expected output:
black yellow pink hills mountain liver
内容的提问来源于stack exchange,提问作者meme
相关产品推荐
相关产品推荐

