如何合并多个DataFrame并解决路径拼接TypeError问题
合并多目录下的jsonl文件为带唯一路径的DataFrame
需求说明
需要处理一个包含多个子文件夹的目录,每个子文件夹内包含imgs文件夹和train.jsonl标签文件。目标是将所有train.jsonl内容合并为一个DataFrame,并将file_name字段更新为父目录路径+图片名称的唯一路径格式。
目录结构
$ tree . ├── sample │ ├---- folder_1 ----|-- -- train.jsonl | | |----- imgs | | | └───├── 0.png | | | └── 1.png | | | └── 2.png | | | └── 3.png .. .. ... ... | | | └── n.png │ ├---- folder_2 ----|-- -- train.jsonl | | |----- imgs | | | └───├── 0.png | | | └── 1.png | | | └── 2.png | | | └── 3.png .. .. ... ... | | | └── n.png │ ├---- folder_3 ----|-- -- train.jsonl | | |----- imgs | | | └───├── 0.png | | | └── 1.png | | | └── 2.png | | | └── 3.png .. .. ... ... | | | └── n.png
示例train.jsonl内容
folder_1中的train.jsonl:
{"file_name": "0.png", "text": "Hello"} {"file_name": "1.png", "text": "there"}
folder_2中的train.jsonl:
{"file_name": "0.png", "text": "Hi"} {"file_name": "1.png", "text": "there from the second dir"}
尝试的代码及报错
尝试使用以下代码合并文件:
import pandas as pd import os df = pd.DataFrame(columns=['file_name', 'text']) # Traverse the directory recursively for root, dirs, files in os.walk('sample'): for file in files: if file == 'train.jsonl': df_temp = pd.read_json(os.path.join(root, file), lines=True) df_temp['file_name'] = os.path.join(root, 'imgs', df_temp['file_name']) df = df.append(df_temp, ignore_index=True) print(df)
运行后报错:
Traceback (most recent call last): File "merage_files.py", line 11, in <module> print(os.path.join(root, 'imgs', df_temp['file_name'])) File "/usr/lib/python3.8/posixpath.py", line 90, in join genericpath._check_arg_types('join', a, *p) File "/usr/lib/python3.8/genericpath.py", line 152, in _check_arg_types raise TypeError(f'{funcname}() argument must be str, bytes, or ' TypeError: join() argument must be str, bytes, or os.PathLike object, not 'Series'
期望结果
最终生成的DataFrame格式如下:
file_name text 0 sample/folder_1/imgs/0.png Hello 1 sample/folder_1/imgs/1.png there 2 sample/folder_2/imgs/0.png Hi 3 sample/folder_2/imgs/1.png there from the second dir .......... ........
解决方案
报错原因是os.path.join只能处理字符串,无法直接接收Pandas Series。需要对Series中的每个元素单独执行路径拼接操作,同时建议用列表收集DataFrame再合并,比append效率更高:
import pandas as pd import os dfs = [] # 遍历目录 for root, _, files in os.walk('sample'): for file in files: if file == 'train.jsonl': # 读取jsonl文件 df_temp = pd.read_json(os.path.join(root, file), lines=True) # 对每个file_name拼接完整路径 df_temp['file_name'] = df_temp['file_name'].apply( lambda x: os.path.join(root, 'imgs', x) ) dfs.append(df_temp) # 合并所有DataFrame final_df = pd.concat(dfs, ignore_index=True) print(final_df)
补充说明
- 使用
apply方法遍历Series中的每个文件名,调用os.path.join生成完整路径 - 用列表
dfs收集所有临时DataFrame,最后用pd.concat合并,避免多次append带来的性能损耗 - 生成的
file_name会包含imgs目录(如sample/folder_1/imgs/0.png),符合实际文件路径结构;如果不需要imgs部分,可修改拼接逻辑为os.path.join(root, x)
内容的提问来源于stack exchange,提问作者Mohammed
相关产品推荐
相关产品推荐

