如何让JavaScript与Python的base64 atob得到相同结果
移植JS的
atob到Python时结果不一致的问题 我知道Python有base64.urlsafe_b64decode(),但想深入理解Base64解码的底层逻辑,所以打算把一段JS实现的atob代码移植到Python,但遇到了结果不一致的问题。
原JS的atob实现代码
function atob (input) { var chars = 'ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/='; var str = String(input).replace(/=+$/, ''); if (str.length % 4 == 1) { throw new InvalidCharacterError("'atob' failed: The string to be decoded is not correctly encoded."); } for ( // initialize result and counters var bc = 0, bs, buffer, idx = 0, output = ''; // get next character buffer = str.charAt(idx++); // character found in table? initialize bit storage and add its ascii value; ~buffer && (bs = bc % 4 ? bs * 64 + buffer : buffer, // and if not first of each 4 characters, // convert the first 8 bits to one ascii character bc++ % 4) ? output += String.fromCharCode(255 & bs >> (-2 * bc & 6)) : 0 ) { // try to find character in table (0-63, not found => -1) buffer = chars.indexOf(buffer); } return output; }
我的困惑点
我看不懂JS中for循环的逻辑,尤其是:
bs = bc % 4 ? bs*64+buffer: buffer, bc++ %4里的逗号表达式String.fromCharCode(255 & bs >> (-2 * bc & 6))的位运算部分
我尝试编写的Python代码
import re # Test subject b64_str: str = "fwHzODWqgMH+NjBq02yeyQ==" # Lookup table for characters chars: str = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+/=" # Replace right padding with empty string replaced = re.sub("=+$", '', b64_str) if len(replaced) % 4 == 1: raise ValueError("atob failed. The string to be decoded is not valid base64") # Bit storage and counters bc = 0 out: str = '' for i in replaced: # Get ascii value of character buffer = ord(i) # If counter is evenly divisible by 4, return buffer as is, else add the ascii value bs = bc * 64 + buffer if bc % 4 else buffer bc += 1 % 4 # Not sure I understand this part # Check if character is in the chars table if i in chars: # Check if the bit storage and bit counter are non-zero if bs and bc: # If so, convert the first 8 bits to an ascii character out += chr(255 & bs >> (-2 * bc & 6)) else: out = 0 # Set buffer to the index of where the first instance of the character is in the b64 string # print(f"before: {chr(buffer)}") buffer = chars.index(chr(buffer)) # print(f"after: {buffer}") print(out)
问题现状
测试字符串"fwHzODWqgMH+NjBq02yeyQ=="的解码结果:
- JS输出:
ó85ªÁþ60jÓlÉ - 我的Python代码输出:
2:u1(²ë:ð1G>%Y
结果不一致,需要解决这个问题。
内容的提问来源于stack exchange,提问作者Coldchain9
相关产品推荐
相关产品推荐

