如何进一步加速从DICOM结构集中提取坐标?已用numpy.reshape优化
Great question! You’re already on the right track using numpy.reshape—that’s one of the biggest wins for this kind of operation. Let’s dive into several more strategies to squeeze out additional speed from your contour parsing workflow:
Key Optimizations to Try
1. Replace len(loop.ContourData)//3 with loop.ContourPointsCount
DICOM’s ContourSequence entries include a ContourPointsCount attribute that directly tells you the number of 3D points. Calculating len(loop.ContourData)//3 adds an unnecessary division and length check; using the pre-existing integer value is faster and more reliable.
2. Minimize Attribute Access Inside Loops
Attribute lookups in Python are relatively slow. Assign frequently accessed attributes to local variables inside your inner function to reduce overhead:
def _reshape_contour_data(loop): data = loop.ContourData count = loop.ContourPointsCount return np.reshape(np.array(data), (3, count))
3. Use np.fromiter Instead of np.array for Faster Conversion
Converting the ContourData list (a sequence of floats) to a NumPy array can be optimized with np.fromiter, which avoids some overhead of np.array for list inputs:
def _reshape_contour_data(loop): count = loop.ContourPointsCount arr = np.fromiter(loop.ContourData, dtype=np.float64, count=3*count) return arr.reshape((3, count))
4. Replace map with List Comprehensions
List comprehensions are often faster than map in Python, especially when combined with local variable optimizations. Rewrite your return statement as:
return [_reshape_contour_data(loop) for loop in contour.ContourSequence]
5. JIT Compile the Inner Function with Numba
Numba can compile your reshaping function to machine code, eliminating Python loop overhead. Here’s how to modify your code:
First install Numba:
pip install numba
Then update your code:
import numba @numba.njit def _reshape_numba(arr, count): return arr.reshape((3, count)) def _reshape_contour_data(loop): count = loop.ContourPointsCount arr = np.fromiter(loop.ContourData, dtype=np.float64, count=3*count) return _reshape_numba(arr, count)
Note: The first call to the Numba-compiled function will have a small compilation delay, but subsequent calls will be much faster.
6. Batch Process All Contours at Once (If Possible)
If all your contours have the same number of points, you can stack all ContourData lists into a single NumPy array and reshape in one go. This is only feasible for datasets with uniform contour lengths, but it’s a huge optimization when applicable:
def parse_coords(contour): if not hasattr(contour, "ContourSequence"): return [] sequences = contour.ContourSequence if not sequences: return [] # Check if all contours have the same point count point_counts = [loop.ContourPointsCount for loop in sequences] if len(set(point_counts)) != 1: # Fall back to per-loop processing return [_reshape_contour_data(loop) for loop in sequences] count = point_counts[0] # Combine all contour data into a single array all_data = np.concatenate([np.fromiter(loop.ContourData, dtype=np.float64) for loop in sequences]) # Reshape to (num_loops, 3, num_points) return all_data.reshape((len(sequences), 3, count))
7. Use PyPy Instead of CPython
If your code doesn’t rely on CPython-specific extensions, running it with PyPy can give significant speedups (often 2-5x) for loop-heavy tasks. Just verify pydicom compatibility (it should work, but test first!).
Profiling to Find Bottlenecks
Don’t forget to keep profiling with cProfile to verify which optimizations are actually moving the needle:
cProfile.runctx('parse_coords(your_contour)', globals(), locals(), 'profile_stats') stats = pstats.Stats('profile_stats') stats.sort_stats(pstats.SortKey.TIME).print_stats(10)
This will show you exactly where your code spends the most time, so you can focus your optimizations effectively.
内容的提问来源于stack exchange,提问作者ColonelFazackerley

