如何保留不同文章目录下重复文件名对应的内容并写入Pandas DataFrame
The issue with your original code is that dictionaries can't have duplicate keys—so when you hit the same filename in a different folder, it overwrites the earlier entry. Instead of using a dictionary, we'll collect each file's data in a list of dictionaries to keep every entry, even when filenames repeat.
Here's the corrected code:
import os import pandas as pd # Initialize an empty list to store all file data data_list = [] root_directory = 'C:/Users/Project/Documents/' # Walk through all directories and files for path, _, files in os.walk(root_directory): for file in files: if file.endswith('.txt'): # Get the full path to the text file full_txt_path = os.path.join(path, file) # Read the text content with open(full_txt_path, 'r') as target_file: text_content = target_file.read() # Generate the corresponding image filename image_filename = file.replace('.txt', '.jpg') # Add the entry to our list data_list.append({ 'image_path': image_filename, 'text': text_content }) # Convert the list to a DataFrame df = pd.DataFrame(data_list) print(df.head())
This will produce exactly the output you're expecting—duplicate image_path values will appear as separate rows with their respective text content.
Optional Enhancement: Add Article Context
If you want to easily distinguish which article each entry comes from, you can add an article column to your DataFrame. Just modify the append step to include the folder name:
data_list.append({ 'article': os.path.basename(path), # Gets the folder name (e.g., "Article1") 'image_path': image_filename, 'text': text_content })
This way you'll have clear visibility into which article each image/text pair belongs to, even when filenames are identical.
内容的提问来源于stack exchange,提问作者Mathew

