You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何合并多个DataFrame并解决路径拼接TypeError问题

合并多目录下的jsonl文件为带唯一路径的DataFrame

需求说明

需要处理一个包含多个子文件夹的目录,每个子文件夹内包含imgs文件夹和train.jsonl标签文件。目标是将所有train.jsonl内容合并为一个DataFrame,并将file_name字段更新为父目录路径+图片名称的唯一路径格式。

目录结构

$ tree
.
├── sample
│   ├---- folder_1 ----|-- -- train.jsonl
|   |                  |----- imgs 
|   |                  |          └───├── 0.png
|   |                  |              └── 1.png
|   |                  |              └── 2.png
|   |                  |              └── 3.png
..  ..                ...               ...
|   |                  |              └── n.png
│   ├---- folder_2 ----|-- -- train.jsonl
|   |                  |----- imgs 
|   |                  |          └───├── 0.png
|   |                  |              └── 1.png
|   |                  |              └── 2.png
|   |                  |              └── 3.png
..  ..                ...               ...
|   |                  |              └── n.png
│   ├---- folder_3 ----|-- -- train.jsonl
|   |                  |----- imgs 
|   |                  |          └───├── 0.png
|   |                  |              └── 1.png
|   |                  |              └── 2.png
|   |                  |              └── 3.png
..  ..                ...               ...
|   |                  |              └── n.png

示例train.jsonl内容

folder_1中的train.jsonl:

{"file_name": "0.png", "text": "Hello"}
{"file_name": "1.png", "text": "there"}

folder_2中的train.jsonl:

{"file_name": "0.png", "text": "Hi"}
{"file_name": "1.png", "text": "there from the second dir"}

尝试的代码及报错

尝试使用以下代码合并文件:

import pandas as pd
import os

df = pd.DataFrame(columns=['file_name', 'text'])

# Traverse the directory recursively
for root, dirs, files in os.walk('sample'):
    for file in files:
        if file == 'train.jsonl':
            df_temp = pd.read_json(os.path.join(root, file), lines=True)
            df_temp['file_name'] = os.path.join(root, 'imgs', df_temp['file_name'])

            df = df.append(df_temp, ignore_index=True)

print(df) 

运行后报错:

Traceback (most recent call last):
  File "merage_files.py", line 11, in <module>
    print(os.path.join(root, 'imgs', df_temp['file_name']))
  File "/usr/lib/python3.8/posixpath.py", line 90, in join
    genericpath._check_arg_types('join', a, *p)
  File "/usr/lib/python3.8/genericpath.py", line 152, in _check_arg_types
    raise TypeError(f'{funcname}() argument must be str, bytes, or '
TypeError: join() argument must be str, bytes, or os.PathLike object, not 'Series'

期望结果

最终生成的DataFrame格式如下:

file_name                    text
0  sample/folder_1/imgs/0.png           Hello
1  sample/folder_1/imgs/1.png           there
2  sample/folder_2/imgs/0.png            Hi
3  sample/folder_2/imgs/1.png         there from the second dir
  ..........                          ........

解决方案

报错原因是os.path.join只能处理字符串,无法直接接收Pandas Series。需要对Series中的每个元素单独执行路径拼接操作,同时建议用列表收集DataFrame再合并,比append效率更高:

import pandas as pd
import os

dfs = []

# 遍历目录
for root, _, files in os.walk('sample'):
    for file in files:
        if file == 'train.jsonl':
            # 读取jsonl文件
            df_temp = pd.read_json(os.path.join(root, file), lines=True)
            # 对每个file_name拼接完整路径
            df_temp['file_name'] = df_temp['file_name'].apply(
                lambda x: os.path.join(root, 'imgs', x)
            )
            dfs.append(df_temp)

# 合并所有DataFrame
final_df = pd.concat(dfs, ignore_index=True)
print(final_df)

补充说明

  • 使用apply方法遍历Series中的每个文件名,调用os.path.join生成完整路径
  • 用列表dfs收集所有临时DataFrame,最后用pd.concat合并,避免多次append带来的性能损耗
  • 生成的file_name会包含imgs目录(如sample/folder_1/imgs/0.png),符合实际文件路径结构;如果不需要imgs部分,可修改拼接逻辑为os.path.join(root, x)

内容的提问来源于stack exchange,提问作者Mohammed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 21:23:11