如何修复Hypothesis生成含Strand列的DataFrame时的排序类型错误?
Let's break down what's going wrong and fix it without converting your Strand strings to integers.
The Root Cause
Your original code worked because you were sorting a tuple of two integers, but once you added the Strand string to the tuple, calling sorted() on the whole thing tries to compare integers and strings—something Python doesn't allow, hence the TypeError: '<' not supported between instances of 'str' and 'int'.
Solution 1: Only Sort the Start/End Values (Simplest Fix)
Instead of sorting the entire row tuple, split it apart, sort just the integer positions, then recombine with the Strand value. This keeps your Strand column as a string and ensures End ≥ Start every time.
Here's the corrected code:
from hypothesis.extra.pandas import columns, data_frames import hypothesis.strategies as st positions = st.integers(min_value=0, max_value=int(1e7)) strands = st.sampled_from("+ -".split()) # Custom function to sort only Start and End, leave Strand untouched def ensure_end_ge_start(row_tuple): start, end, strand = row_tuple sorted_start_end = sorted([start, end]) return (sorted_start_end[0], sorted_start_end[1], strand) # Also fix the dtype parameter—don't force all columns to be int! df = data_frames( columns=columns( ["Start", "End", "Strand"], dtype={"Start": int, "End": int, "Strand": str} ), rows=st.tuples(positions, positions, strands).map(ensure_end_ge_start) ).example()
Key notes here:
- We explicitly split the tuple into its components, sort only the two integers, then put everything back together. No cross-type comparisons happen here.
- We updated the
dtypeargument to a dictionary so we don't accidentally force the Strand column to be an integer (that was a hidden bug waiting to bite even if the sorting worked!).
Solution 2: Pre-Generate Ordered Start/End Pairs (More "Hypothesis-idiomatic")
If you prefer to build the ordering into your strategy instead of using a map() step, you can create a strategy that generates (Start, End) pairs where Start ≤ End right from the start, then combine it with your Strand strategy.
Example code:
from hypothesis.extra.pandas import columns, data_frames import hypothesis.strategies as st # Create a strategy that generates ordered (Start, End) pairs directly ordered_positions = st.builds( lambda a, b: (min(a, b), max(a, b)), st.integers(min_value=0, max_value=int(1e7)), st.integers(min_value=0, max_value=int(1e7)) ) strands = st.sampled_from("+ -".split()) # Combine the ordered positions with Strand to make row tuples row_strategy = ordered_positions.flatmap( lambda pos: st.tuples(st.just(pos[0]), st.just(pos[1]), strands) ) df = data_frames( columns=columns( ["Start", "End", "Strand"], dtype={"Start": int, "End": int, "Strand": str} ), rows=row_strategy ).example()
This approach keeps all the logic within Hypothesis's strategy system, which can be cleaner for more complex data generation setups.
Either solution will fix your TypeError while keeping the Strand column as a string, which is exactly what you wanted.
内容的提问来源于stack exchange,提问作者The Unfun Cat

