You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python逐行排序文件(禁用列表/全加载)时出现乱码问题求助

Fixing Garbled Text & Sorting Logic in Your SortedFile Implementation

Hey there! Let's work through the issues in your create function—both the garbled text problem and the broken sorting logic, since your assignment requires proper line-by-line sorting without loading the whole file into memory or using lists.

What's Causing the Garbled Text & Broken Sorting?

  1. File Pointer Chaos with w+ Mode: When you open your target file with w+, after writing the first line, the file pointer sits at the end of the file. When you try to loop over heapfile right after, you're reading from the end—so you get no content at all. This leads to messy, unordered writes that look like garbled text.
  2. Flawed findIndex Logic: Your findIndex checks if the file is empty to return -1, but when you first create the file, the header hasn't been written yet. This would lead to an index error when trying to access line.split(',')[i] later.
  3. Incomplete Sorting Logic: The current code doesn't actually implement line-by-line insertion sort. You're not properly comparing lines and inserting new content in the correct position.

Corrected Implementation

Here's a revised version that follows your assignment rules (line-by-line processing, no lists, allowed temporary files):

import os

class SortedFile:
    def __init__(self, file_name, col_name):
        self.name = file_name
        self.sort_col = col_name
        # Initialize empty file
        with open(self.name, "w") as f:
            pass

    def is_empty(self):
        with open(self.name, "r") as heapfile:
            return heapfile.readline() == ''

    def _get_col_index(self, header_line):
        # Get index from header instead of hardcoding (more robust)
        cols = header_line.strip().split(',')
        try:
            return cols.index(self.sort_col)
        except ValueError:
            raise ValueError(f"Column '{self.sort_col}' not found in header")

    def create(self, source_file):
        with open(source_file, "r") as f:
            # Read header first
            header = f.readline()
            if not header:
                return  # Empty source file
            
            # Write header to sorted file first
            with open(self.name, "w") as sorted_file:
                sorted_file.write(header)
            
            # Get the index of the column we're sorting by
            sort_idx = self._get_col_index(header)

            # Process each data line one by one
            for line in f:
                line = line.strip()
                if not line:
                    continue  # Skip empty lines
                
                # Extract the value we'll compare for sorting
                new_val = line.split(',')[sort_idx]
                temp_file = f"{self.name}.tmp"

                # Open sorted file for reading, temp file for writing
                with open(self.name, "r") as sorted_f, open(temp_file, "w") as temp_f:
                    # Write header to temp file first
                    temp_f.write(header)
                    inserted = False

                    # Iterate through existing sorted lines to find insertion point
                    for sorted_line in sorted_f:
                        sorted_line_stripped = sorted_line.strip()
                        if not sorted_line_stripped:
                            continue
                        
                        # Skip header line (we already wrote it)
                        if sorted_line_stripped == header.strip():
                            continue
                        
                        sorted_val = sorted_line_stripped.split(',')[sort_idx]
                        # Insert the new line before the first larger sorted value
                        if sorted_val > new_val:
                            temp_f.write(line + '\n')
                            inserted = True
                        # Write the existing sorted line
                        temp_f.write(sorted_line)
                    
                    # If we reached the end without inserting, add the new line at the end
                    if not inserted:
                        temp_f.write(line + '\n')
                
                # Replace the sorted file with the temp file
                os.replace(temp_file, self.name)

# Usage example
if __name__ == "__main__":
    sf = SortedFile('SortedFile.txt', 'currency')
    sf.create('kiva.txt')

Key Changes Explained

  • File Handling with Temp Files: For each new line, we use a temporary file to build the new sorted content. This avoids loading the entire file into memory and ensures we insert each line in the correct position.
  • Header-Based Column Index: Instead of hardcoding column indices, we extract the index from the source file's header—this makes the code more flexible if columns are reordered.
  • Proper Insertion Logic: We iterate through the existing sorted lines, compare the sort values, and insert the new line exactly where it belongs. If it's the largest value so far, we add it to the end.
  • Clean File Operations: Using with statements ensures files are properly closed, and os.replace safely swaps the temp file with the main sorted file.

Why This Fixes Garbled Text

The original code's w+ mode and incorrect pointer handling led to partial writes and empty reads. By using separate read/write operations with temporary files, we ensure every line is written in order without overwriting or skipping content.

内容的提问来源于stack exchange,提问作者Itzik Dan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:30:56