如何在不使用unsafe代码的情况下将SHA1哈希info_hash字节数组转换为BitTorrent Tracker公告URL可用的编码字符串?
Great question! The core issue here is that your original unsafe code treats arbitrary binary data (the SHA1 hash) as valid UTF-8 string data—which it's not. That's why you see garbled output when printing h, even though the Url crate quietly handles the raw bytes correctly under the hood. Let's fix this safely, no unsafe code required.
Using the urlencoding Crate (Simplest & Recommended Approach)
The easiest way to handle URL encoding for binary data is to use the dedicated urlencoding crate, which is designed for exactly this use case and avoids unsafe operations entirely.
- First, add the dependency to your
Cargo.toml:
[dependencies] url = "2.4" urlencoding = "2.1"
- Update your code to encode the byte array directly:
use url::Url; use urlencoding::encode_binary; fn main() { let h: [u8; 20] = [216, 247, 57, 206, 195, 40, 149, 108, 204, 91, 191, 31, 134, 217, 253, 207, 219, 168, 206, 182]; // Encode the binary hash to the URL-safe string format let encoded_hash = encode_binary(&h).to_string(); println!("{}", encoded_hash); // Outputs: %D8%F79%CE%C3%28%95l%CC%5B%BF%1F%86%D9%FD%CF%DB%A8%CE%B6 let url = Url::parse_with_params("http://bttracker.org:6969/test", &[("info_hash", &encoded_hash)]).unwrap().to_string(); println!("{}", url); // Same correct URL as before // This assertion will now pass! assert_eq!(encoded_hash, "%D8%F79%CE%C3%28%95l%CC%5B%BF%1F%86%D9%FD%CF%DB%A8%CE%B6"); }
Manual Implementation (For Understanding the Encoding Process)
If you want to see how URL encoding works under the hood, you can implement it manually by iterating over each byte and converting non-unreserved characters to the %XX format:
use url::Url; fn url_encode_binary(data: &[u8]) -> String { // Preallocate space to avoid unnecessary reallocations (each byte needs up to 3 characters) let mut result = String::with_capacity(data.len() * 3); for &byte in data { match byte { // Unreserved characters (letters, numbers, and a few symbols) don't need encoding b'a'..=b'z' | b'A'..=b'Z' | b'0'..=b'9' | b'-' | b'_' | b'.' | b'~' => { result.push(byte as char); } // All other bytes get converted to %XX hex format _ => { result.push_str(&format!("%{:02X}", byte)); } } } result } fn main() { let h: [u8; 20] = [216, 247, 57, 206, 195, 40, 149, 108, 204, 91, 191, 31, 134, 217, 253, 207, 219, 168, 206, 182]; let encoded_hash = url_encode_binary(&h); println!("{}", encoded_hash); // Same correct output as before assert_eq!(encoded_hash, "%D8%F79%CE%C3%28%95l%CC%5B%BF%1F%86%D9%FD%CF%DB%A8%CE%B6"); }
Why Your Original Unsafe Code Was Problematic
std::str::from_utf8_uncheckedforces the compiler to treat arbitrary bytes as valid UTF-8, but SHA1 hashes contain bytes that are not valid UTF-8 sequences. This is undefined behavior—while it didn't crash immediately, it's unsafe and could lead to bugs or crashes in more complex code.- The garbled print output happens because your terminal tries to interpret the invalid UTF-8 bytes as text, which fails. The
Urlcrate worked around this because it processes the raw byte data directly (instead of treating it as a UTF-8 string) when encoding parameters.
内容的提问来源于stack exchange,提问作者Kevin

