如何实现含禁用字符字符串的可逆转换以适配特殊存储?
& and =) Got it, let's break down how to solve this problem. The core needs here are 100% reversibility, no ambiguity, and complete removal of the forbidden characters & and =. Below are two practical, easy-to-implement solutions depending on your use case:
Solution 1: Custom Escape Mapping (Readable, Text-Friendly)
This approach is ideal if you want the encoded string to remain mostly human-readable. We'll create unique, non-conflicting escape sequences for the forbidden chars, and first escape the escape marker itself to avoid collisions.
Encoding Steps:
- Escape the escape marker first: Replace all instances of
__(our chosen marker prefix) with____—this ensures existing__in your input doesn't get confused with our escape sequences. - Replace forbidden chars:
- Swap
&with__AMP__ - Swap
=with__EQ__
- Swap
Decoding Steps (Reverse Order):
- Restore forbidden chars:
- Swap
__AMP__back to& - Swap
__EQ__back to=
- Swap
- Restore original marker: Replace
____back to__
Example Code (Python):
def encode_forbidden_chars(s: str) -> str: # Step 1: Escape the double underscore marker s = s.replace("__", "____") # Step 2: Replace forbidden characters s = s.replace("&", "__AMP__") s = s.replace("=", "__EQ__") return s def decode_forbidden_chars(s: str) -> str: # Step 1: Restore forbidden characters s = s.replace("__AMP__", "&") s = s.replace("__EQ__", "=") # Step 2: Restore original double underscores s = s.replace("____", "__") return s # Test it out original = "Hello & World = __test__" encoded = encode_forbidden_chars(original) decoded = decode_forbidden_chars(encoded) print(f"Original: {original}") print(f"Encoded: {encoded}") # Output: Hello __AMP__ World __EQ__ ____test____ print(f"Decoded: {decoded}") # Output matches original
Solution 2: Modified Base64 (Universal, for Arbitrary Data)
If you're dealing with non-text data or want a solution that handles all special chars (not just & and =), a modified Base64 approach works great. Standard Base64 only uses = for padding, which is forbidden here—we'll replace that padding with a safe character like -.
Encoding Steps:
- Convert your input string to bytes (using UTF-8 for text).
- Encode the bytes with standard Base64.
- Replace all
=padding characters with-(or any other safe char not in your forbidden list).
Decoding Steps:
- Replace
-back to=to restore the original Base64 format. - Decode the Base64 string back to bytes.
- Convert bytes back to a UTF-8 string.
Example Code (Python):
import base64 def encode_base64_safe(s: str) -> str: bytes_data = s.encode("utf-8") base64_str = base64.b64encode(bytes_data).decode("utf-8") # Replace forbidden = with - return base64_str.replace("=", "-") def decode_base64_safe(s: str) -> str: # Restore padding base64_str = s.replace("-", "=") bytes_data = base64.b64decode(base64_str) return bytes_data.decode("utf-8") # Test it out original = "UserID=123&Name=Alice" encoded = encode_base64_safe(original) decoded = decode_base64_safe(encoded) print(f"Original: {original}") print(f"Encoded: {encoded}") # Output: VXNlcklEPTEyMyZOYW1lPUFsaWNl- (varies slightly) print(f"Decoded: {decoded}") # Output matches original
Which One to Choose?
- Use the custom escape mapping if you need the encoded string to be readable (e.g., log messages, user-facing text).
- Use the modified Base64 if you're handling binary data, or want a one-size-fits-all solution for all special characters.
内容的提问来源于stack exchange,提问作者AlikElzin-kilaka

