PySpark中reduceByKey()报TypeError: 'float'不可迭代错误求助
Hey Helen, let's figure out why you're hitting that TypeError: 'float' object is not iterable with reduceByKey() in your Spark code.
First, let's break down the error: The stack trace points straight to your lambda function in the reduceByKey call. This happens when your combiner function (that lambda) tries to iterate over a float value—and since floats aren't iterable (unlike lists or tuples), it throws this error.
Looking at your field_valid filter, you're keeping entries where dis, TxP, ef, and pl aren't NaN. I assume after filtering, you're mapping your data into key-value pairs where the value is a float. The problem boils down to how your reduceByKey lambda is handling that float value.
Common Scenarios & Fixes
Let's cover the most likely cases based on your error:
You're trying to sum values per key but used the wrong lambda
If you wrote something likereduceByKey(lambda a,b: sum(a) + sum(b)), that's the issue. The first time the lambda runs,ais a single float—callingsum(a)tries to iterate over that float, which triggers the error.The fix is straightforward: just add the floats directly, no
sum()needed:# Correct way to sum floats per key summed_rdd = your_filtered_rdd.reduceByKey(lambda a, b: a + b)You're trying to collect all values per key into a list
If you triedreduceByKey(lambda a,b: a + [b]), the firstais a float, and you can't add a float to a list. Instead, use eithergroupByKey(simpler for collecting values) oraggregateByKey(more efficient for large datasets):# Option 1: Use groupByKey to collect values into a list value_list_rdd = your_filtered_rdd.groupByKey().mapValues(list) # Option 2: Use aggregateByKey (better performance for large data) value_list_rdd = your_filtered_rdd.aggregateByKey( [], # Initial empty list for each key lambda acc, val: acc.append(val) or acc, # Add new value to the list (append returns None, so we use `or acc` to return the updated list) lambda acc1, acc2: acc1 + acc2 # Merge lists from different partitions )
Quick Troubleshooting Tip
Go back to your reduceByKey line and inspect the lambda function. Ask yourself: Is this lambda trying to iterate over the input values (which are floats)? If yes, adjust the logic to handle single float values instead of treating them as iterables.
For example, if you accidentally used a function that expects an iterable (like sum() on a single float), replace that with direct arithmetic or list operations that work with individual floats.
内容的提问来源于stack exchange,提问作者Helen Z

