Python脚本批量文件流式预处理需求及Shell语法错误咨询
First, let's tackle that pesky syntax error, then move on to your batch processing goal.
Resolving the -c: line 0: syntax error near unexpected token '('' Error
This error almost always comes from how you're invoking your Python script in the shell—specifically, unescaped special characters (like parentheses, apostrophes, or spaces) in your input/output filenames, or incorrect quoting when passing arguments. Here's how to fix it:
- Quote filenames with special characters: Wrap any filename containing spaces, parentheses, apostrophes, or other shell metacharacters in double quotes. For example:
python your_script.py "my input (file).txt" "output file.txt" - Escape apostrophes if using single quotes: If you prefer single quotes, escape apostrophes with a backslash:
python your_script.py 'file\'s name.txt' output.txt - Check for command substitution issues: If you're using something like
python your_script.py $(generate_filename) output.txt, make sure the output ofgenerate_filenameis properly escaped. Useprintf "%q"to safely handle special characters:python your_script.py "$(generate_filename | xargs printf "%q")" output.txt - Double-check your command line: Ensure you're not mixing up quotes or adding extra parentheses by mistake. The error means the shell is misparsing your command before it even reaches Python.
Processing Files Without Saving Preprocessed Versions
Since you have 1000+ files and want to skip saving intermediate preprocessed files, here are two practical approaches depending on where your preprocessing logic lives:
Option 1: Preprocess in Memory Within Your Python Script
Modify your script to read the input file, perform preprocessing directly in memory, then use the processed data for your main logic. This keeps everything in RAM until you write the final output. Here's a simplified example:
import sys def preprocess_content(content): # Replace this with your actual preprocessing steps processed = content.replace("old_text", "new_text").strip() return processed def main(): if len(sys.argv) != 3: print("Usage: python your_script.py <input_file> <output_file>") sys.exit(1) input_path = sys.argv[1] output_path = sys.argv[2] # Read input, preprocess in memory with open(input_path, 'r') as f: raw_content = f.read() processed_content = preprocess_content(raw_content) # Run your main script logic on processed_content # ... (your existing code here, adjusted to use processed_content instead of reading from a file) # Write final output with open(output_path, 'w') as f: f.write(processed_content) if __name__ == "__main__": main()
Option 2: Pipe Preprocessed Content to Your Script (If Using Shell Tools)
If your preprocessing steps use shell commands like sed or awk, modify your script to read from standard input (stdin) when you pass - as the input file. Then you can pipe preprocessed content directly into your script without saving intermediate files.
First, update your script to handle stdin:
import sys def preprocess_content(content): # Your preprocessing logic here (or skip if using shell tools) return content def main(): # Read from stdin if input is '-' if len(sys.argv) >=2 and sys.argv[1] == '-': raw_content = sys.stdin.read() output_path = sys.argv[2] else: if len(sys.argv) !=3: print("Usage: python your_script.py <input_file/-> <output_file>") sys.exit(1) input_path = sys.argv[1] output_path = sys.argv[2] with open(input_path, 'r') as f: raw_content = f.read() processed_content = preprocess_content(raw_content) # Main logic and output writing... with open(output_path, 'w') as f: f.write(processed_content) if __name__ == "__main__": main()
Then use a shell loop to process all files:
for input_file in /path/to/your/files/*.txt; do # Replace with your actual shell preprocessing command sed 's/old_pattern/new_pattern/g' "$input_file" | python your_script.py - "processed_${input_file}" done
This pipeline skips writing any intermediate files entirely, saving storage and processing time.
内容的提问来源于stack exchange,提问作者rex

