You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

输入图像尺寸/分辨率对跨数据集语义分割网络输出质量的影响

How Input Resolution Impacts Semantic Segmentation Performance When Transferring Between Datasets

Great question—this is an extremely common challenge when moving semantic segmentation models from a training dataset like Cityscapes to a test dataset like KITTI, where native resolutions don’t align (Cityscapes typically uses 1024x2048, while KITTI’s raw images sit around 375x1242). Let’s break down exactly how input size affects your model’s output quality, and why you’re seeing gaps between subjective perception and (m)IoU scores.

Key Impacts of Mismatched Resolution

1. Misaligned Receptive Fields

Semantic segmentation models (think FCN, DeepLab, or U-Net) learn a receptive field—the area of the input image each neuron "sees"—tuned directly to the training dataset’s resolution. When you feed in a smaller KITTI image:

  • The model’s receptive field covers a smaller physical area of the scene, so it can’t capture the context needed to classify large objects (like buildings or trucks) correctly.
  • Smaller objects (like pedestrians or cyclists) in KITTI occupy far fewer pixels than they do in Cityscapes, making it hard for the model to detect and segment them consistently.

2. Feature Map Scale Errors

Most segmentation models rely on repeated downsampling (pooling) and upsampling operations optimized for the training input size. When you change the input resolution:

  • Downsampling steps may produce non-integer feature map sizes, forcing interpolation that introduces noise or blurs fine-grained details.
  • Upsampling layers (like transposed convolutions) are calibrated to reconstruct features from training-scale feature maps. Mismatched input sizes lead to misaligned boundaries and distorted object shapes, which hits subjective quality hard.

3. Dataset-Specific Object Scale Bias

Cityscapes and KITTI don’t just differ in resolution—they have distinct scene dynamics that amplify the problem:

  • Cityscapes focuses on dense urban environments where objects are relatively large and evenly distributed.
  • KITTI’s ego-vehicle perspective means objects shrink rapidly with distance, creating a wider range of object scales that the model never learned to handle during training.
    This scale mismatch leads to both low (m)IoU (from missed or misclassified small objects) and poor subjective quality (from blurry or misaligned large objects).

Why Subjective Quality and (m)IoU Diverge

  • Subjective perception prioritizes sharp boundaries, complete object masks, and visually coherent scene segmentation. Even small errors in boundary alignment or object completeness make results look "bad" to the human eye.
  • (m)IoU is a pixel-level metric that averages performance across all classes. Poor segmentation of small, numerous classes (like pedestrians) can drag down the overall score, even if large objects look passable. Conversely, blurry boundaries might not hit (m)IoU as hard as full object misclassifications, but they’re far more noticeable visually.

Practical Fixes to Mitigate the Issue

  • Preserve aspect ratio when resizing: Instead of stretching KITTI images to fit Cityscapes’ 1024x2048, use letterbox resizing (add black bars to maintain the original 3.31:1 aspect ratio) or crop to a compatible aspect ratio. This reduces shape distortion that confuses the model.
  • Multi-scale training/testing: If you can re-train or fine-tune, add random resizing to your training pipeline (e.g., resize inputs to 800x1600, 1024x2048, 1280x2560) to teach the model to handle varying resolutions. For testing, run inference on multiple scales and average the results to boost robustness.
  • Fine-tune on KITTI: Use your pre-trained Cityscapes model as a starting point and fine-tune it on KITTI’s dataset. This lets the model learn KITTI’s specific object scales, resolution, and scene context. Start with freezing the model’s backbone (to retain general features) and only train the segmentation head, then gradually unfreeze layers for better adaptation.
  • Adjust receptive fields: Replace standard convolutions with dilated convolutions in the model’s backbone. Dilated convolutions expand the receptive field without increasing the number of parameters, helping the model capture more context from smaller input images.

内容的提问来源于stack exchange,提问作者crossx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 08:10:35