如何在Python中从命令行读取多个输入文件并实现词频统计
Handling Multiple Input Files in Python for Word Frequency Count
Hey there! To adapt your script to handle multiple files passed via the command line, you just need to loop through all the file paths provided after the script name. Here's how you can do it:
Modified Code for Combined Word Counts
This version aggregates word frequencies from all input files into a single dictionary, giving you a total count across all documents:
import sys def count_words_in_file(file_path): """Helper function to count words in a single file""" word_count = {} with open(file_path, 'r') as file: for line in file: # Split line into words, trimming whitespace first words = line.strip().split() for word in words: # Use dict.get() to simplify count updates word_count[word] = word_count.get(word, 0) + 1 return word_count def main(): # Check if any files were provided if len(sys.argv) < 2: print("Usage: python script.py file1.txt file2.txt ... fileN.txt") sys.exit(1) combined_counts = {} # Loop through all provided file paths (skip sys.argv[0] which is the script name) for file_path in sys.argv[1:]: try: file_counts = count_words_in_file(file_path) # Merge individual file counts into the combined dictionary for word, count in file_counts.items(): combined_counts[word] = combined_counts.get(word, 0) + count except FileNotFoundError: print(f"Warning: File '{file_path}' not found, skipping.") except Exception as e: print(f"Error processing '{file_path}': {str(e)}") # Print sorted results for readability print("Combined Word Frequencies:") for word, count in sorted(combined_counts.items()): print(f"{word}: {count}") if __name__ == "__main__": main()
Key Changes Explained
- Loop through all input files: We use
sys.argv[1:]to capture every argument after the script name—these are your target file paths. - Helper function: Extracting the word-counting logic makes the code cleaner and easier to maintain.
- Error handling: Added checks for missing files and unexpected errors so the script keeps running even if one file is problematic.
- Simplified count updates: Using
dict.get()eliminates the need for conditional checks to see if a word already exists in the dictionary.
Alternative: Separate Counts per File
If you want to keep word counts isolated for each file instead of combining them, modify the main function like this:
def main(): if len(sys.argv) < 2: print("Usage: python script.py file1.txt file2.txt ... fileN.txt") sys.exit(1) # Store counts in a dictionary where keys are file paths file_word_counts = {} for file_path in sys.argv[1:]: try: file_word_counts[file_path] = count_words_in_file(file_path) except FileNotFoundError: print(f"Warning: File '{file_path}' not found, skipping.") except Exception as e: print(f"Error processing '{file_path}': {str(e)}") # Print results for each file for file_path, counts in file_word_counts.items(): print(f"\nWord Frequencies for {file_path}:") for word, count in sorted(counts.items()): print(f"{word}: {count}")
How to Run
Just pass your files as command-line arguments exactly like you wanted:
python script.py file1.txt file2.txt file3.txt
This setup will handle as many files as you need to process! 😊
内容的提问来源于stack exchange,提问作者Al-Amin
相关产品推荐
相关产品推荐

