如何使用Python读取文件夹文本并按对应JPG图片名保存文件,以及为批量关联图片生成同名ground-truth文件
Hey there! Let's work through your two Python batch processing tasks step by step. I'll break down the logic for each one and share practical, commented code snippets you can use right away.
需求一:读取文本内容并生成对应JPG同名TXT文件
核心思路
Assuming your setup is:
- A folder with your JPG images (
img_folder) - A folder with corresponding text files (either with the same base name as the JPGs, e.g.,
photo1.jpg↔photo1.txt, or mapped by matching prefixes) - A target folder where you want to save the final TXT files
The steps are:
- Iterate over all JPG files in the image folder
- For each JPG, find its matching text file in the source text folder
- Read the text content
- Save the content as a TXT file with the same base name as the JPG in the target folder
Python Code
from pathlib import Path def jpg_to_txt(img_folder: str, text_folder: str, target_folder: str): # Convert paths to Path objects for cleaner handling img_path = Path(img_folder) text_path = Path(text_folder) target_path = Path(target_folder) # Create target folder if it doesn't exist target_path.mkdir(parents=True, exist_ok=True) # Iterate over all JPG files in the image folder for img_file in img_path.glob("*.jpg"): # Get the base name of the image (without .jpg extension) img_base = img_file.stem # Find the corresponding text file (assuming same base name, .txt extension) text_file = text_path / f"{img_base}.txt" if text_file.exists(): # Read the text content with open(text_file, "r", encoding="utf-8") as f: content = f.read() # Save to target folder with same base name as JPG target_file = target_path / f"{img_base}.txt" with open(target_file, "w", encoding="utf-8") as f: f.write(content) print(f"Saved: {target_file}") else: print(f"Warning: No text file found for {img_file.name}") # Example usage jpg_to_txt( img_folder="./your_images", text_folder="./your_texts", target_folder="./output_txt" )
Key Notes
- If your text files don't have exactly the same base name as the JPGs, adjust the
text_filepath logic (e.g., use prefix matching withglob). pathlibeliminates messy string concatenation for paths—way easier to maintain!encoding="utf-8"ensures we handle special characters correctly.
需求二:批量生成关联图片的Ground-Truth文件
核心思路
From your example, here's the pattern we need to handle:
- Ground-truth file:
[prefix].gt.txt(e.g.,Doc0006.Row1City.gt.txt) - Associated images:
[prefix].gt,[prefix]0_rotate.jpg,[prefix]1_rotate.jpg, etc.
The steps are:
- Iterate over all ground-truth files (
*.gt.txt) in your source folder - Extract the core prefix (e.g.,
Doc0006.Row1CityfromDoc0006.Row1City.gt.txt) - Find all image files that start with this prefix
- For each image, create a TXT file with the same name as the image, using the content from the ground-truth file
- Save all these TXT files to your target folder
Python Code
from pathlib import Path def generate_gt_files(source_folder: str, target_folder: str): source_path = Path(source_folder) target_path = Path(target_folder) # Create target folder if it doesn't exist target_path.mkdir(parents=True, exist_ok=True) # Iterate over all ground-truth files for gt_file in source_path.glob("*.gt.txt"): # Extract the core prefix (remove .gt from the stem) gt_prefix = gt_file.stem.replace(".gt", "") gt_content = gt_file.read_text(encoding="utf-8") # Find all files starting with the core prefix (skip the GT file itself) for image_file in source_path.glob(f"{gt_prefix}*"): if image_file == gt_file: continue # Generate target TXT file with same name as the image txt_filename = f"{image_file.stem}.txt" target_file = target_path / txt_filename # Write ground-truth content to the target file target_file.write_text(gt_content, encoding="utf-8") print(f"Generated GT file: {target_file}") # Example usage generate_gt_files( source_folder="./your_source_files", target_folder="./output_gt" )
Key Notes
- The
gt_file.stem.replace(".gt", "")trick perfectly extracts the core prefix from your example naming pattern. glob(f"{gt_prefix}*")catches all related files, including.gtfiles and*_rotate.jpgfiles.- If your image naming varies, you can refine the match with regex (e.g.,
re.match(rf"{gt_prefix}\d*_rotate\.jpg", image_file.name)to only target rotated JPGs).
内容的提问来源于stack exchange,提问作者Faheem-Ur-Rehman 2305-FETBSEEF
相关产品推荐
相关产品推荐

