PySpark如何将PipelinedRDD转换为分号分隔的坐标字符串
Got it, let's break down how to turn your PipelinedRDD of coordinate tuples into that specific semicolon-separated string format for URL parameters. Here's a straightforward, practical approach:
Step 1: Format individual coordinate pairs
First, use map() to transform each (longitude, latitude) tuple in your RDD into a comma-separated string. This handles the inner formatting for every coordinate pair:
# Transform each tuple to "lon,lat" string formatted_coords_rdd = all_coord_iso_rdd.map(lambda coord: f"{coord[0]},{coord[1]}")
Step 2: Collect and join into a single string
Next, collect the formatted strings from the RDD to your driver node (note: this works best if your dataset is small enough for URL parameters—since URLs have strict length limits, this is usually safe here), then join all the strings with semicolons:
# Join all formatted coordinate strings with semicolons final_param_string = ";".join(formatted_coords_rdd.collect())
Example Output
Using your sample data from all_coord_iso_rdd.take(4), this will produce exactly the string you're targeting:
"-73.57534790039062,45.5311393737793;-73.574951171875,45.529457092285156;-73.5749282836914,45.52922821044922;-73.57501220703125,45.52901077270508"
Quick Heads-Up
If you were working with an extremely large RDD, collect() could cause memory issues on the driver. But since this is for URL parameters, you're almost certainly dealing with a manageable number of coordinates—so this method should work flawlessly.
内容的提问来源于stack exchange,提问作者adil blanco

