You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现从文件路径提取含数字的文件夹名或文件名

路径解析解决方案:提取含数字的文件夹/文件名

需求规则

  • 优先提取路径中纯数字命名的文件夹,返回该数字
  • 若路径中无纯数字文件夹,提取文件名中包含的数字,返回该数字
  • 以上都不满足时,返回文件名作为默认值

目录结构示例

project
   |_62951
        |_test1.docx
   |_68512
        |_test2.docx
        |_minor tasks
             |_test3.docx
   |_Plumbing project
        |_69251
             |_test4.dox
        |_House address
             |_69251 plumb.docx
             |_test5.docx

预期输出

project code: 62951
filename: test1.docx
project code: 68512
filename: test2.docx
project code: 68512
filename: test3.docx
project code: 69251
filename: test4.docx
project code: 69251
filename: 69251 plumb.docx
project code: test5.docx
filename: test5.docx

现有代码

import os

#run through all folders
def get_files(source):
    matches = []
    for root, dirnames, filenames in os.walk(source):
        for filename in filenames:
                matches.append(os.path.join(root, filename))
    return matches


def parse(files):
    
    # run through all files
    folders = []
    for file in files:
        filepath,filename = os.path.split(file)
        filebreak = [filepath.split("\\")]
        print('project code: %s' % filebreak)
        print('file name: %s' % filename)


        #check file name

path = 'C:\\Users\\quan.nguyen\\***\\***\\Project testing files\\XML'

parse(get_files(path))

当前问题

现有代码仅拆分了路径并输出文件夹列表,未按需求规则检查纯数字文件夹、提取文件名中的数字,无法得到预期结果。

解决方案代码

import os
import re

def get_files(source):
    matches = []
    for root, dirnames, filenames in os.walk(source):
        for filename in filenames:
            matches.append(os.path.join(root, filename))
    return matches


def parse(files):
    for file in files:
        filepath, filename = os.path.split(file)
        # 拆分路径为所有文件夹组成的列表,适配不同系统路径分隔符
        folders = filepath.split(os.sep)
        
        project_code = None
        # 遍历所有文件夹,查找纯数字命名的文件夹
        for folder in folders:
            if folder.isdigit():
                project_code = folder
                break  # 找到第一个纯数字文件夹即停止
        
        # 若未找到纯数字文件夹,检查文件名中的数字
        if not project_code:
            # 提取文件名中的所有数字序列,取第一个匹配项
            num_matches = re.findall(r'\d+', filename)
            if num_matches:
                project_code = num_matches[0]
            else:
                # 都不满足则返回文件名作为默认值
                project_code = filename
        
        print(f'project code: {project_code}')
        print(f'filename: {filename}')


path = 'C:\\Users\\quan.nguyen\\***\\***\\Project testing files\\XML'
parse(get_files(path))

关键说明

  1. 跨平台路径拆分:使用os.sep代替硬编码的\\,适配Windows、Linux等不同操作系统的路径分隔符
  2. 纯数字文件夹检查:遍历路径中的每个文件夹,通过str.isdigit()判断是否为纯数字命名,找到即返回
  3. 文件名数字提取:用正则表达式r'\d+'匹配文件名中的数字序列,取第一个匹配结果
  4. 默认值处理:当以上条件都不满足时,直接返回文件名作为项目代码

内容的提问来源于stack exchange,提问作者Quan Hoang Nguyen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 22:30:31