Python处理公式时eval莫名报语法错误,请求技术排查
Hey there, let's figure out why you're hitting that SyntaxError at the 501st formula for the first sample. I've dealt with similar pandas parsing quirks before, so here's a breakdown of the likely issues and how to fix them:
Key Problem: Incorrectly Extracting Formula Text from pandas
Your current code converts an entire pandas Series (df.iloc[j]) to a string, which includes extra metadata like index labels, dtype info, and sometimes truncated text for long formulas. That's why your formula looks complete in the file but gets chopped off in code.
Let's Break Down the Issues in Your Current Code
str(df.iloc[j])doesn't give you just the formula text—it returns something like:"0 a+b*c m=0.56\nName: 0, dtype: object"- When formulas are long, pandas automatically truncates the string with
...to fit display limits, so you end up with an incomplete formula foreval(). - Your logic to strip extra content (
f[1:].strip(), splitting on\n) is fragile and depends on pandas' string formatting, which can vary between rows.
Step-by-Step Fixes
1. First: Validate Your Formula DataFrame Structure
Before fixing code, confirm you're reading the formula file correctly. Print the raw value of the problematic formula (index 500, since we start at 0) to see its true content:
print(repr(df.iloc[500])) # Use repr() to show hidden characters/truncation
This will tell you if pandas is storing the full formula, or if there are unexpected newlines/characters in the file.
2. Extract Formula Text Properly
Instead of converting the entire Series to a string, directly access the cell's raw text. Assuming you loaded your formulas into a single-column DataFrame (e.g., named formula), use:
formula_raw = df.iloc[j]['formula'] # Get the pure text of the formula
If you didn't name the column, use df.iloc[j].iloc[0] to grab the first (and only) value in the row.
3. Clean the Formula Safely
Strip the m=xxx suffix reliably using split() instead of relying on newlines:
# Split on " m=" and take the first part (the actual formula) formula_clean = formula_raw.split(' m=')[0].strip()
This works regardless of what follows m= and avoids issues with line breaks in formulas.
4. Add Error Handling to Diagnose Issues
Wrap your evaluation in a try-except block to capture exactly which formula is failing and what its content is:
for i, r in data.iterrows(): vars_dict = r.to_dict() # Store sample variables in a dict (safer for eval) for j in range(len(df)): try: formula_raw = df.iloc[j]['formula'] formula_clean = formula_raw.split(' m=')[0].strip() result = eval(formula_clean, {}, vars_dict) # Pass variables safely print(result, end=" ") except SyntaxError as se: print(f"\n--- Error at Sample {i+1}, Formula {j+1} ---") print(f"Raw formula: {repr(formula_raw)}") print(f"Cleaned formula sent to eval: {repr(formula_clean)}") raise se # Stop execution to debug, or comment out to continue print("\n")
5. Bonus: Safer Variable Handling
Using eval(formula_clean, {}, vars_dict) is better than assigning variables to the global namespace—it keeps your variables isolated and avoids accidental overwrites.
Why This Fixes the Truncation Issue
By accessing the raw cell text instead of converting the Series to a string, you bypass pandas' automatic truncation of long strings. This ensures you're always sending the full, unmodified formula to eval().
Important Note About Security
eval() can execute arbitrary code, so never use this if your formula file comes from an untrusted source. For safer formula parsing, consider using libraries like sympy to parse and evaluate mathematical expressions without the security risks.
内容的提问来源于stack exchange,提问作者mitra

