PySpark遍历列替换标点为空格报错:Column对象不可调用
Hey, I see the issue here—your function name pd is causing a name collision, which is why it worked for one column but broke when you added the second.
Chances are you've got a line somewhere in your code like import pandas as pd (super common alias!) and when you try to call your pd() function, Python is mixing it up with the pandas alias (or maybe even a PySpark Column object that ended up named pd). That's exactly why you're getting the "Column object is not callable" error—it's trying to treat a Column like a function.
Here's the easy fix: just rename your function to something that doesn't clash with common aliases, like process_text_cols:
from pyspark.sql import functions as F def process_text_cols(data): df = data df = df.select('oproblem', 'lca') text_col = ['oproblem', 'lca'] for i in text_col: df = df.withColumn(i, F.lower(F.col(i))) df = df.withColumn(i, F.regexp_replace(F.col(i), '[.,#-:;/?!\']', ' ')) return df
What's going on here?
- The single-column version worked by luck—maybe the name conflict didn't trigger in that specific execution flow. But once you added the second column and expanded the loop, the interpreter resolved
pdto a Column object instead of your custom function, leading to the error. - By using a unique function name, you eliminate the collision entirely.
One quick check to confirm: scan your code for any imports that use pd as an alias (like pandas or PySpark pandas utilities). If you need to keep those imports, just make sure your text processing function has a distinct name and you'll be good to go.
内容的提问来源于stack exchange,提问作者PineNuts0

