Cryptopals挑战1第6题:重复密钥XOR破解Keysize计算错误求助
Fixing Keysize Detection in Cryptopals Set 1 Challenge 6
Let's walk through fixing your keysize detection issue step by step. I've spotted a few critical problems in your code that are leading to the incorrect keysize (20 instead of 29), and we'll address each one:
Key Issues in Your Original Code
- Variable Name Typo: You used
for key in range(2,41):but referencedkeysizelater in the loop—this would throw aNameErrorunless it's a typo in your pasted code. It should befor keysize in range(2,41):. - Skipping Base64 Decoding: The challenge file
6.txtis Base64-encoded. You're processing the raw Base64 string instead of decoding it to the actual encrypted byte stream, which completely invalidates your Hamming distance calculations. - String vs. Byte Handling: Your Hamming function uses
ord()on string characters, but we need to work directly with bytes for accurate distance calculations (since the encryption operates on bytes). - Suboptimal Keysize Scoring: You're processing block pairs sequentially and averaging all of them together, but the challenge recommends using multiple blocks (like the first 4) and averaging their pairwise distances for a more stable score.
Corrected Python 3.7.6 Code
Here's the fixed version with explanations:
import base64 def hamming_distance(b1: bytes, b2: bytes) -> int: """Calculate Hamming distance between two byte strings (matches challenge test case).""" return sum(bin(byte1 ^ byte2).count('1') for byte1, byte2 in zip(b1, b2)) # Test the Hamming function (verifies it returns 37 as required) t1 = b"this is a test" t2 = b"wokka wokka!!!" assert hamming_distance(t1, t2) == 37, "Hamming distance function is incorrect" # Read and decode the Base64 encrypted file with open('6.txt', 'rb') as f: ciphertext = base64.b64decode(f.read().strip()) def find_best_keysize(ciphertext: bytes, min_keysize: int = 2, max_keysize: int = 40) -> int: scores = {} for keysize in range(min_keysize, max_keysize + 1): # Grab the first 4 blocks of the current keysize (per challenge suggestion) blocks = [ciphertext[i*keysize : (i+1)*keysize] for i in range(4)] pairwise_distances = [] # Calculate distance between all unique block pairs for i in range(len(blocks)): for j in range(i + 1, len(blocks)): # Skip if blocks are uneven (only happens at the very end of the ciphertext) if len(blocks[i]) != len(blocks[j]): continue dist = hamming_distance(blocks[i], blocks[j]) normalized_dist = dist / keysize # Normalize by keysize pairwise_distances.append(normalized_dist) # Average the distances for this keysize scores[keysize] = sum(pairwise_distances) / len(pairwise_distances) # Return the keysize with the smallest average normalized distance return min(scores, key=scores.get) best_keysize = find_best_keysize(ciphertext) print(f"Detected best keysize: {best_keysize}") # This will output 29
What Changed and Why
- Base64 Decoding: We read the file as bytes and decode it using
base64.b64decode()—this gives us the actual encrypted byte stream that the challenge expects us to work with. - Byte-Oriented Hamming Function: The function now takes byte strings directly, which eliminates any issues with character encoding and matches how XOR encryption works at the byte level.
- Improved Keysize Scoring: We use the first 4 blocks for each keysize, calculate all pairwise Hamming distances, normalize each by keysize, then average those values. This aligns with the challenge's recommendation and produces a more reliable score for each keysize.
- Test Assertion: Added an assertion to confirm the Hamming function works correctly (critical for debugging).
When you run this code, it should correctly detect the keysize as 29, which is the expected result for the challenge.
内容的提问来源于stack exchange,提问作者eli
相关产品推荐
相关产品推荐

