You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Spark 2.3:子目录读取的通配符问题及通用实现方案

Spark TextFile Subdirectory Reading: Wildcards & Generic Paths

Great question — let’s break this down clearly, since handling nested directories in Spark is a common gotcha.

Do I have to use wildcards to read subdirectories?

Short answer: No, but the default behavior makes it seem that way.

By default, Spark’s textFile() only reads direct files in the specified directory — it won’t automatically traverse into subdirectories. That’s why your test showed wildcards were necessary to access files in /Users/test/data/folder1/.

That said, you can avoid wildcards entirely by enabling Spark’s recursive file lookup feature (available in Spark 2.0+). We’ll cover that in the next section.

How to handle variable/deeper subdirectory levels?

If your directory structure has unknown or varying levels of subdirectories, you have two reliable solutions:

1. Use the /** wildcard for recursive matching

The /** wildcard is designed to match all levels of subdirectories and files under a parent path. This is the simplest way to write a generic path that works regardless of how many subdirectory layers exist.

Example code:

String folder = "/Users/test/data/**";
sparkContext.textFile(folder, 1).toJavaRDD();

This will pick up every file under /Users/test/data/, whether it’s in /data/file.txt, /data/folder1/file.txt, /data/folder1/subfolder2/file.txt, etc.

2. Enable recursive file lookup via configuration

If you prefer not to use wildcards, you can set a Spark configuration property to enable automatic recursive traversal. This lets you just specify the parent directory path, and Spark will handle the rest.

Example setup:

// Configure Spark to enable recursive file lookup
SparkConf conf = new SparkConf()
    .setAppName("RecursiveFileReader")
    .setMaster("local") // Adjust for your cluster environment
    .set("spark.sql.sources.recursiveFileLookup.enabled", "true");

SparkContext sparkContext = new SparkContext(conf);

// Now just specify the parent directory
String folder = "/Users/test/data";
sparkContext.textFile(folder, 1).toJavaRDD();

Bonus: wholeTextFiles() for path-aware reading

If you need to track which file each line comes from, you can use wholeTextFiles() instead. This method automatically recursively reads subdirectories and returns a JavaPairRDD<String, String> where the key is the full file path, and the value is the entire content of the file. Just note this is different from textFile() which splits files into individual lines.


内容的提问来源于stack exchange,提问作者Molay

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:42:22