You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于Shadow Detection论文中图像平均Patch数量的计算疑问

Hey there, let's unpack why you're seeing such a huge discrepancy between your patch count and what the paper reports! Here are the key reasons and actionable checks to resolve this:

Possible Reasons for the Patch Count Mismatch

1. You're Extracting All Possible Patches, But the Paper Uses Random Sampling

The biggest culprit here is likely that you're generating every possible patch via sliding windows with tf.extract_image_patches, but the paper randomly samples only ~13 total instances per image across the three categories.

GAN training often avoids using every single patch from an image (which leads to redundant data and massive compute overhead). Instead, authors typically sample a small, fixed number of patches per image to keep training efficient. The paper's "average 13 training instances/patch" refers to the total number of sampled patches per image across all three types, not the total number of possible patches you can extract with sliding windows.

2. Misinterpreting the "3/4 of the original image" Patch Size

The paper mentions patches of "3/4 of the original image"—it’s easy to assume this means 75% of the input image’s height/width, but there’s a chance it refers to 75% of the 256×256 target size (i.e., 192×192 patches) extracted from the original image, then downsampled to 256×256.

If you use a fixed 192×192 patch size instead of scaling with the original image’s dimensions, the number of possible patches drops significantly. Even then, the paper would only sample a subset of these patches to hit the 13-instance average.

3. Unaccounted Dataset Preprocessing Steps

The SBU Shadow Dataset has a wide range of image sizes, but the paper might have applied preprocessing you’re missing:

  • They could have resized all images to a fixed maximum dimension (e.g., 400px on the longest side) before patch extraction, reducing the total number of possible patches.
  • They might have filtered out images with extreme aspect ratios or smaller dimensions, which could pull down the average patch count per image.

Next Steps to Verify

  • Re-read the paper’s 3.4 section carefully for any mentions of "random sampling" or "subset selection"—many papers gloss over this detail but it’s critical to their training setup.
  • Try sampling 13 total instances per image (e.g., 1 resized image, 6 patches from category 2, 6 from category 3) and see if this aligns with the paper’s intended training pipeline.
  • Double-check your tf.extract_image_patches parameters: ensure you’re using padding='VALID' (the default) to only extract patches fully contained within the image, as SAME padding would generate even more patches.

内容的提问来源于stack exchange,提问作者wadhwasahil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 04:13:17