PySpark中map与list相加触发TypeError的技术求助
Hey there, I see exactly what's going on here! That error pops up because in Python 3, the map() function returns an iterator object, not a list—and you can't use the + operator to combine an iterator with a list directly.
The Quick Fix
Wrap your map() call with list() to convert that iterator into a proper list, which can then be safely concatenated with your numerical columns list:
# Instead of this (which causes the error) # featureColumns = map(lambda c: c + "classVec", categoricalColumns) + numericalColumns # Do this: convert map to list first encodedCatColumns = list(map(lambda c: c + "classVec", categoricalColumns)) featureColumns = encodedCatColumns + numericalColumns # Then use featureColumns in VectorAssembler assembler = VectorAssembler(inputCols=featureColumns, outputCol="features")
A More Pythonic Alternative
If you prefer, you can use a list comprehension instead of map()—it's more readable and returns a list directly, so no extra conversion needed:
featureColumns = [c + "classVec" for c in categoricalColumns] + numericalColumns
Why This Works
The + operator in Python only works when both operands are of the same sequence type (like two lists). Since map() gives you an iterator (not a list), Python doesn't know how to combine it with your numerical columns list. Converting the iterator to a list (either with list() or a comprehension) makes the types compatible, so the concatenation works smoothly.
Give that a try, and your VectorAssembler should run without that TypeError!
内容的提问来源于stack exchange,提问作者Jabernet

