Python多进程运行脚本报错:'NoneType'对象无'geq'属性求助
Hey there! Let's break down your problem step by step and work through both the error cause and the sequential execution requirement.
First: Diagnose the AttributeError: 'NoneType' object has no attribute 'geq'
That error is a clue pointing to one core issue: somewhere in your failing script, the object you're trying to use the >= operator on is None. Here's why that happens in your func1 code:
f=df.where((df[a] >= dat1) & (df[b] <= dat2))
- The
geqmentioned in the error is actually pandas' underlying method for the>=operator (__ge__). The error pops up because either:- Your
dfitself isNone: The script failed to load the DataFrame correctly (e.g., wrong file path, missing file, or read error that wasn't caught). - Columns
aorbdon't exist indf: If you reference a column that isn't present, pandas might returnNone(or throw aKeyError—but if you have a try/except swallowing that error, it could leavedf[a]/df[b]asNone). dat1ordat2isNone: Comparing a Series toNonecan lead to unexpected behavior, though this usually returns a boolean Series rather than aNoneTypeerror.
- Your
Quick test to narrow it down: Run the failing script outside of multiprocessing (just execute it directly). If it still throws the same error, the problem is in the script itself, not the multiprocessing setup.
Second: Run Scripts Sequentially Without Early Exit
Right now, if you're starting both processes at once or not handling failures properly, your main script might exit early. To fix this, you need to:
- Start one process at a time.
- Wait for each process to finish (success or failure) before launching the next.
- Add error handling in both the child processes and main process to avoid crashes.
Example Fixed Code
from multiprocessing import Process import pandas as pd def run_script(path): try: # Load DataFrame (use absolute path to avoid working directory issues in subprocesses) df = pd.read_csv("/full/path/to/your/data.csv") # Execute script-specific logic if path == "1.py": # Insert func1 and other 1.py logic here a = "valid_column_a" b = "valid_column_b" dat1 = 10 # Replace with your actual value dat2 = 50 # Replace with your actual value f = df.where((df[a] >= dat1) & (df[b] <= dat2)) print(f"Script {path} ran successfully") elif path == "2.py": # Insert 2.py logic here print(f"Script {path} ran successfully") except Exception as e: print(f"Script {path} failed: {str(e)}") if __name__ == "__main__": scripts = ["1.py", "2.py"] # Run scripts one after another for script in scripts: p = Process(target=run_script, args=(script,)) p.start() p.join() # Wait for the current script to finish before moving to the next
Key Fixes in This Code
- Sequential execution: Using
join()ensures the main process waits for each subprocess to complete (whether it succeeds or fails) before starting the next one. - Absolute file paths: Subprocesses might have a different working directory than your main script—using absolute paths for your DataFile avoids "file not found" errors that could leave
dfasNone. - Error handling: Wrapping the script logic in
try/exceptcatches errors and prints them without crashing the subprocess or main process.
Final Troubleshooting Steps
- Verify the failing script works standalone: Run it directly to confirm the DataFrame loads correctly and columns
a/bexist. - Check
dat1/dat2values: Ensure they're notNoneand are compatible with the data type of columnsa/b(e.g., don't compare a numeric column to a string). - Avoid sharing DataFrames across processes: If you were passing the main process's
dfto subprocesses, this can cause serialization issues. Let each subprocess load its own copy of the DataFrame instead.
内容的提问来源于stack exchange,提问作者NewBee

