如何解码HTML文件中的Base64编码字符串并替换为解码结果?
Got it, let's solve this step by step. You're dealing with Base64 encoded chunks in your HTML that follow the pattern [any content]base64,[encoded string], and you need to decode those chunks and replace the original strings with the decoded content. Here's how to do it using Python's regex and base64 modules, plus integration with BeautifulSoup if you're using that.
Step 1: Regex Pattern & Decoding Function
We'll use a regex to capture the encoded Base64 part, then a helper function to decode it and replace the original match. The regex is designed to non-greedily match everything up to base64,, then capture valid Base64 characters (A-Z, a-z, 0-9, +, /, and = for padding).
Here's the code:
import re import base64 def decode_base64_match(match): # Capture the prefix (everything before "base64,") and the encoded string prefix, encoded = match.groups() try: # Decode the Base64 bytes (strip whitespace to handle accidental line breaks) decoded_bytes = base64.b64decode(encoded.strip()) # Convert bytes to UTF-8 string (adjust encoding if your content uses something else) decoded_str = decoded_bytes.decode('utf-8') # Replace the entire original string with the decoded content return decoded_str except Exception as e: # If decoding fails (invalid Base64), leave the original string and log a warning print(f"Warning: Failed to decode chunk starting with {encoded[:20]}...: {str(e)}") return match.group(0)
Step 2: Process HTML as a Raw String
If you're treating the entire HTML as a single string, apply the regex substitution directly:
# Example input HTML original_html = """ <style> charset=utf-8;base64,I2JhY2tydW5uZXJfUV81c3R7aGVpZ2h0OjkzcHg7fWJhY2tydW5uZXJfUV81c3R7ZGlzcGxheTpibG9jayFpbXBvcnRhbnQ7fQ== </style> """ # Process the HTML processed_html = re.sub( r'(.*?base64,)([A-Za-z0-9+/=\s]+)', decode_base64_match, original_html, flags=re.DOTALL # Allow matching across lines if needed ) print(processed_html)
Running this will replace the Base64 chunk with the decoded CSS:
<style> #backrunner_Q_5st{height:93px;}backrunner_Q_5st{display:block!important;} </style>
Step 3: Integrate with BeautifulSoup
If you're using BeautifulSoup to parse HTML, you can apply the same logic to specific elements or all text nodes:
Process Specific Elements (e.g., Style Tags)
from bs4 import BeautifulSoup soup = BeautifulSoup(original_html, 'html.parser') # Target all <style> tags and process their content for style_tag in soup.find_all('style'): if style_tag.string: style_tag.string = re.sub( r'(.*?base64,)([A-Za-z0-9+/=\s]+)', decode_base64_match, style_tag.string ) # Get the processed HTML processed_html_with_soup = str(soup)
Process All Text Nodes
If Base64 chunks could be anywhere in text content, iterate over all text nodes:
for text_node in soup.find_all(text=True): new_text = re.sub( r'(.*?base64,)([A-Za-z0-9+/=\s]+)', decode_base64_match, text_node ) text_node.replace_with(new_text) processed_html = str(soup)
Key Notes
- Encoding: The example uses UTF-8, but adjust the
decode()call if your content uses a different encoding (like ISO-8859-1). - Whitespace: The regex includes
\sto handle accidental line breaks or spaces in the Base64 string—base64.b64decode ignores whitespace automatically. - Error Handling: The try/except block ensures your code doesn't crash on invalid Base64; it just leaves the original string intact and logs a warning.
内容的提问来源于stack exchange,提问作者lgjmac

