基于JDK8的Groovy实现UTF-8转EUC-JP并生成MD5哈希求助
UTF-8 to EUC-JP Conversion + MD5 Hash in Groovy (JDK8)
Got it, let's fix this for you. The key issue with your initial attempt was unnecessary intermediate String conversions that can corrupt the byte values we need for the hash. Let's break down exactly what your PHP code does, then replicate that behavior perfectly in Groovy.
What your PHP code does under the hood:
- Takes a UTF-8 encoded string (as a raw byte sequence).
- Converts those bytes directly to EUC-JP encoded bytes.
- Computes the MD5 hash of those EUC-JP bytes and outputs it as a lowercase hex string.
The Groovy Solution:
We need to work directly with byte arrays for encoding and hashing to avoid data loss. Here's the complete working code:
import java.security.MessageDigest import java.nio.charset.Charset def getMd5AfterEucJpConversion(String utf8Input) { // Step 1: Encode the input string (Unicode sequence) to EUC-JP bytes Charset eucJpCharset = Charset.forName("EUC-JP") byte[] eucJpBytes = utf8Input.getBytes(eucJpCharset) // Step 2: Compute MD5 hash of the raw EUC-JP bytes MessageDigest md5Digest = MessageDigest.getInstance("MD5") byte[] hashBytes = md5Digest.digest(eucJpBytes) // Step 3: Convert hash bytes to lowercase hex string (matches PHP's md5() output) StringBuilder hexResult = new StringBuilder() hashBytes.each { byte b -> String hex = Integer.toHexString(0xFF & b) // Pad single-digit hex values with a leading zero to match PHP's format if (hex.length() == 1) hexResult.append('0') hexResult.append(hex) } return hexResult.toString() } // Example usage def testInput = "こんにちは" // Sample UTF-8 string def md5Output = getMd5AfterEucJpConversion(testInput) println(md5Output) // This will match your PHP code's output for the same input
Why this works:
- No data corruption: We skip unnecessary String-to-byte conversions and directly encode the input (which represents your original UTF-8 content as Unicode characters) to EUC-JP bytes.
- Exact MD5 match: We compute the hash on the raw EUC-JP bytes just like PHP does. The final hex conversion ensures we get the same lowercase, zero-padded format as PHP's
md5()function.
Important Notes:
- Ensure your input String in Groovy is properly decoded from UTF-8 bytes (e.g., use
new File("file.txt").getText("UTF-8")when reading files). - EUC-JP doesn't support all Unicode characters—unsupported characters will be replaced with a default (usually
?), matching PHP's default behavior formb_convert_encoding.
内容的提问来源于stack exchange,提问作者Md. Hasan Basri
相关产品推荐
相关产品推荐

