使用glob遍历指定目录时输出顺序与输入一致是否可靠?
Great question! The short answer is: No, this consistency isn’t guaranteed to hold in all cases, but in your specific current setup, it’s working because of how you’re pairing glob with your fixed codes list. Let’s break this down clearly.
Why It’s Working Right Now
Looking at your code, you’re looping through each entry in your explicitly ordered codes list one by one. For each code, you run:
directories = glob.glob("../../DATA/{}".format(code))
Since your pattern is an exact directory name (not a wildcard like *), glob will either return an empty list (if the directory doesn’t exist) or a list with exactly one path. When there’s only one match, there’s no "order" to debate—you’re just grabbing that single entry, processing it, and appending its resolution to your list in the same sequence as your codes loop.
This is why your output matches the codes order right now: the loop over your predefined list is driving the sequence of results, not glob’s internal behavior.
Why It’s Not a Reliable Long-Term Assumption
The consistency you’re seeing could break if any of these scenarios occur:
- Multiple matches for a single code: If you ever end up with multiple directories matching a code pattern (e.g., accidental duplicates or looser wildcards),
globwill return all matches—but their order depends on your operating system’s file system implementation, not yourcodeslist. This would immediately disrupt your result sequence. - Filesystem-specific quirks: Even with single matches, while most modern file systems return single entries predictably, the
globmodule’s official documentation explicitly states that results are not sorted. There’s no guarantee this behavior won’t shift across Python versions or operating systems.
A More Reliable Approach
Since you already have a fixed order in codes, you don’t need to rely on glob to locate these paths at all. Instead, construct the paths directly and check if they exist. This eliminates any uncertainty from glob entirely.
Here’s a cleaner, more robust rewrite using pathlib (Python 3.4+):
from pathlib import Path import yaml codes = ['a1','b1','c1', 'a2','b2','c2'] resolutions = [] for code in codes: dir_path = Path("../../DATA") / code yaml_path = dir_path / f"{code}.yaml" if yaml_path.exists(): with open(yaml_path) as stream: yaml_content = yaml.load(stream, Loader=yaml.FullLoader) resolutions.append(yaml_content["Resolution"]) else: # Handle missing files/directories as needed resolutions.append(None) # Or skip, log an error, etc. print(resolutions)
This approach uses your codes order to build paths directly, so your result list will always match the sequence you want—no surprises from glob or filesystem quirks.
内容的提问来源于stack exchange,提问作者HungryMolecule

