You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python读取文件夹文本并按对应JPG图片名保存文件,以及为批量关联图片生成同名ground-truth文件

Hey there! Let's work through your two Python batch processing tasks step by step. I'll break down the logic for each one and share practical, commented code snippets you can use right away.

需求一:读取文本内容并生成对应JPG同名TXT文件

核心思路

Assuming your setup is:

  • A folder with your JPG images (img_folder)
  • A folder with corresponding text files (either with the same base name as the JPGs, e.g., photo1.jpg ↔ photo1.txt, or mapped by matching prefixes)
  • A target folder where you want to save the final TXT files

The steps are:

  1. Iterate over all JPG files in the image folder
  2. For each JPG, find its matching text file in the source text folder
  3. Read the text content
  4. Save the content as a TXT file with the same base name as the JPG in the target folder

Python Code

from pathlib import Path

def jpg_to_txt(img_folder: str, text_folder: str, target_folder: str):
    # Convert paths to Path objects for cleaner handling
    img_path = Path(img_folder)
    text_path = Path(text_folder)
    target_path = Path(target_folder)
    
    # Create target folder if it doesn't exist
    target_path.mkdir(parents=True, exist_ok=True)
    
    # Iterate over all JPG files in the image folder
    for img_file in img_path.glob("*.jpg"):
        # Get the base name of the image (without .jpg extension)
        img_base = img_file.stem
        
        # Find the corresponding text file (assuming same base name, .txt extension)
        text_file = text_path / f"{img_base}.txt"
        
        if text_file.exists():
            # Read the text content
            with open(text_file, "r", encoding="utf-8") as f:
                content = f.read()
            
            # Save to target folder with same base name as JPG
            target_file = target_path / f"{img_base}.txt"
            with open(target_file, "w", encoding="utf-8") as f:
                f.write(content)
            
            print(f"Saved: {target_file}")
        else:
            print(f"Warning: No text file found for {img_file.name}")

# Example usage
jpg_to_txt(
    img_folder="./your_images",
    text_folder="./your_texts",
    target_folder="./output_txt"
)

Key Notes

  • If your text files don't have exactly the same base name as the JPGs, adjust the text_file path logic (e.g., use prefix matching with glob).
  • pathlib eliminates messy string concatenation for paths—way easier to maintain!
  • encoding="utf-8" ensures we handle special characters correctly.
需求二:批量生成关联图片的Ground-Truth文件

核心思路

From your example, here's the pattern we need to handle:

  • Ground-truth file: [prefix].gt.txt (e.g., Doc0006.Row1City.gt.txt)
  • Associated images: [prefix].gt, [prefix]0_rotate.jpg, [prefix]1_rotate.jpg, etc.

The steps are:

  1. Iterate over all ground-truth files (*.gt.txt) in your source folder
  2. Extract the core prefix (e.g., Doc0006.Row1City from Doc0006.Row1City.gt.txt)
  3. Find all image files that start with this prefix
  4. For each image, create a TXT file with the same name as the image, using the content from the ground-truth file
  5. Save all these TXT files to your target folder

Python Code

from pathlib import Path

def generate_gt_files(source_folder: str, target_folder: str):
    source_path = Path(source_folder)
    target_path = Path(target_folder)
    
    # Create target folder if it doesn't exist
    target_path.mkdir(parents=True, exist_ok=True)
    
    # Iterate over all ground-truth files
    for gt_file in source_path.glob("*.gt.txt"):
        # Extract the core prefix (remove .gt from the stem)
        gt_prefix = gt_file.stem.replace(".gt", "")
        gt_content = gt_file.read_text(encoding="utf-8")
        
        # Find all files starting with the core prefix (skip the GT file itself)
        for image_file in source_path.glob(f"{gt_prefix}*"):
            if image_file == gt_file:
                continue
            
            # Generate target TXT file with same name as the image
            txt_filename = f"{image_file.stem}.txt"
            target_file = target_path / txt_filename
            
            # Write ground-truth content to the target file
            target_file.write_text(gt_content, encoding="utf-8")
            print(f"Generated GT file: {target_file}")

# Example usage
generate_gt_files(
    source_folder="./your_source_files",
    target_folder="./output_gt"
)

Key Notes

  • The gt_file.stem.replace(".gt", "") trick perfectly extracts the core prefix from your example naming pattern.
  • glob(f"{gt_prefix}*") catches all related files, including .gt files and *_rotate.jpg files.
  • If your image naming varies, you can refine the match with regex (e.g., re.match(rf"{gt_prefix}\d*_rotate\.jpg", image_file.name) to only target rotated JPGs).

内容的提问来源于stack exchange,提问作者Faheem-Ur-Rehman 2305-FETBSEEF

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 05:12:42