Python逐行排序文件(禁用列表/全加载)时出现乱码问题求助
Fixing Garbled Text & Sorting Logic in Your SortedFile Implementation
Hey there! Let's work through the issues in your create function—both the garbled text problem and the broken sorting logic, since your assignment requires proper line-by-line sorting without loading the whole file into memory or using lists.
What's Causing the Garbled Text & Broken Sorting?
- File Pointer Chaos with
w+Mode: When you open your target file withw+, after writing the first line, the file pointer sits at the end of the file. When you try to loop overheapfileright after, you're reading from the end—so you get no content at all. This leads to messy, unordered writes that look like garbled text. - Flawed
findIndexLogic: YourfindIndexchecks if the file is empty to return -1, but when you first create the file, the header hasn't been written yet. This would lead to an index error when trying to accessline.split(',')[i]later. - Incomplete Sorting Logic: The current code doesn't actually implement line-by-line insertion sort. You're not properly comparing lines and inserting new content in the correct position.
Corrected Implementation
Here's a revised version that follows your assignment rules (line-by-line processing, no lists, allowed temporary files):
import os class SortedFile: def __init__(self, file_name, col_name): self.name = file_name self.sort_col = col_name # Initialize empty file with open(self.name, "w") as f: pass def is_empty(self): with open(self.name, "r") as heapfile: return heapfile.readline() == '' def _get_col_index(self, header_line): # Get index from header instead of hardcoding (more robust) cols = header_line.strip().split(',') try: return cols.index(self.sort_col) except ValueError: raise ValueError(f"Column '{self.sort_col}' not found in header") def create(self, source_file): with open(source_file, "r") as f: # Read header first header = f.readline() if not header: return # Empty source file # Write header to sorted file first with open(self.name, "w") as sorted_file: sorted_file.write(header) # Get the index of the column we're sorting by sort_idx = self._get_col_index(header) # Process each data line one by one for line in f: line = line.strip() if not line: continue # Skip empty lines # Extract the value we'll compare for sorting new_val = line.split(',')[sort_idx] temp_file = f"{self.name}.tmp" # Open sorted file for reading, temp file for writing with open(self.name, "r") as sorted_f, open(temp_file, "w") as temp_f: # Write header to temp file first temp_f.write(header) inserted = False # Iterate through existing sorted lines to find insertion point for sorted_line in sorted_f: sorted_line_stripped = sorted_line.strip() if not sorted_line_stripped: continue # Skip header line (we already wrote it) if sorted_line_stripped == header.strip(): continue sorted_val = sorted_line_stripped.split(',')[sort_idx] # Insert the new line before the first larger sorted value if sorted_val > new_val: temp_f.write(line + '\n') inserted = True # Write the existing sorted line temp_f.write(sorted_line) # If we reached the end without inserting, add the new line at the end if not inserted: temp_f.write(line + '\n') # Replace the sorted file with the temp file os.replace(temp_file, self.name) # Usage example if __name__ == "__main__": sf = SortedFile('SortedFile.txt', 'currency') sf.create('kiva.txt')
Key Changes Explained
- File Handling with Temp Files: For each new line, we use a temporary file to build the new sorted content. This avoids loading the entire file into memory and ensures we insert each line in the correct position.
- Header-Based Column Index: Instead of hardcoding column indices, we extract the index from the source file's header—this makes the code more flexible if columns are reordered.
- Proper Insertion Logic: We iterate through the existing sorted lines, compare the sort values, and insert the new line exactly where it belongs. If it's the largest value so far, we add it to the end.
- Clean File Operations: Using
withstatements ensures files are properly closed, andos.replacesafely swaps the temp file with the main sorted file.
Why This Fixes Garbled Text
The original code's w+ mode and incorrect pointer handling led to partial writes and empty reads. By using separate read/write operations with temporary files, we ensure every line is written in order without overwriting or skipping content.
内容的提问来源于stack exchange,提问作者Itzik Dan
相关产品推荐
相关产品推荐

