如何使用spark.read方法排除文件首行(及末行)?
在Databricks中用PySpark读取文件时排除首尾行的解决方案
你的目标文件内容如下:
return (<X> example </X> )ID
11111
22222return (<X> example </X>)
之前尝试的header='false'参数无效,因为这个参数是用来控制是否将首行识别为表头,不是跳过首行;而修改大文件添加注释的方式会增加IO开销,确实没必要。下面提供几种高效的解决方案:
方法一:按行号精准过滤(DataFrame方式)
先读取所有行并添加连续行号,再过滤掉首尾行:
# 读取所有文本行 text_df = spark.read.text("path/to/your/file.txt") # 添加连续行号索引(从0开始) from pyspark.sql.window import Window from pyspark.sql.functions import row_number window_spec = Window.orderBy("value") indexed_df = text_df.withColumn("index", row_number().over(window_spec) - 1) # 获取总行数 total_rows = indexed_df.count() # 过滤首尾行并移除索引列 filtered_df = indexed_df.filter((indexed_df.index != 0) & (indexed_df.index != total_rows - 1)).drop("index")
方法二:RDD方式(适合大文件,行号更精准)
RDD的zipWithIndex能生成严格连续的行号,适合处理大文件:
# 读取文件为RDD text_rdd = spark.sparkContext.textFile("path/to/your/file.txt") # 获取总行数 total_rows = text_rdd.count() # 添加行号并过滤首尾行,再转成DataFrame filtered_rdd = text_rdd.zipWithIndex().filter(lambda x: x[1] != 0 and x[1] != total_rows - 1).map(lambda x: x[0]) filtered_df = filtered_rdd.toDF("value")
方法三:按固定内容过滤(最简单,适合你的场景)
如果首尾行内容固定为<X> example </X>,直接过滤掉包含该内容的行即可:
df = spark.read.text("path/to/your/file.txt") filtered_df = df.filter(~df.value.contains("<X> example </X>"))
内容的提问来源于stack exchange,提问作者HABLOH
相关产品推荐
相关产品推荐

