Windows Media Foundation中WriteSample报MF_E_NO_SAMPLE_DURATION的技术问询
Media Foundation SinkWriter H.264 RTP to MP4 Questions
Great questions about Media Foundation's SinkWriter behavior when handling H.264 RTP streams and exporting to MP4! Let's break down each one with practical context:
1. Why is setting Sample Duration mandatory?
You’d think SinkWriter could calculate duration from consecutive sample times, but the reality ties directly to how Media Foundation and MP4 file structure work:
- MP4 container requirements: MP4 files store per-frame timing information explicitly in their
mdatandmoovatoms. SinkWriter needs this duration to correctly build the media timeline, calculate bitrate estimates, and ensure the file plays back with proper pacing. - SinkWriter’s design constraints: It doesn’t assume you’ll provide consecutive samples (e.g., you might drop frames, add out-of-order samples, or end the stream with a single frame). Relying on future samples to calculate duration would force it to cache indefinitely, which isn’t feasible for real-time scenarios.
- Edge cases: For the last frame in the stream, there’s no "next sample" to compute duration from—so you have to provide it explicitly anyway.
2. Why does WriteSample only throw MF_E_NO_SAMPLE_DURATION after ~1.48 seconds?
This boils down to SinkWriter’s internal buffering strategy:
- SinkWriter doesn’t validate every sample immediately when you call
WriteSample. It caches a batch of samples first to optimize I/O, process H.264 GOP (Group of Pictures) structures, or build a coherent timeline segment. - The 1.48-second threshold is likely tied to its default buffer size or GOP processing logic. It might wait until it has a full GOP or a specific duration of media before attempting to finalize timing metadata. Once it tries to process that buffered batch and finds missing durations, it throws the error.
- It’s not about a specific frame—it’s about the point when the internal buffer reaches a threshold where timing validation becomes necessary.
3. How to set Sample Duration with uneven frame intervals (15fps average)?
Let’s break down both approaches and their tradeoffs:
3.1 Cache frames to calculate duration from the next sample's timestamp
This is the most accurate and standards-compliant approach:
- For each frame except the last one, cache it until the next frame arrives. Then set the current frame’s duration as
next_sample_time - current_sample_time. - For the final frame, you can either use the average duration of previous frames or estimate it based on the last few intervals.
- Pros: Produces a perfectly accurate media timeline, critical for use cases like video editing, syncing with audio, or precise playback timing.
- Cons: Adds a small amount of latency (equal to one frame interval) since you have to wait for the next frame before writing the current one.
3.2 Use average duration based on frame rate
This is a practical, low-latency workaround that works well for most scenarios:
- For 15fps, the average duration is
10,000,000 / 15 ≈ 666,667(in Media Foundation’s 100-nanosecond time units). - As you’ve tested, this generates playable MP4 files with no visual issues because most media players are tolerant of minor timing discrepancies. The player will adjust playback pacing to match the average rate, and small variations in frame intervals won’t be noticeable to viewers.
- Is this "compliant"? Strictly speaking, MP4 allows variable frame rates, but using average duration doesn’t violate core specs—it just simplifies the timing data. It’s widely accepted for real-time recording scenarios where latency is more important than microsecond-perfect timing.
Which one should you choose?
- Pick 3.1 if you need precise timing (e.g., syncing with audio, professional video production).
- Pick 3.2 if you want low latency and frame interval variations are small enough that viewers won’t notice (e.g., surveillance footage, casual streaming).
内容的提问来源于stack exchange,提问作者karthik vr
相关产品推荐
相关产品推荐

