Python技术问询:如何移除或替换列表中的冗余字符与内容以构建合法字典
Hey there! Let's work through cleaning up that messy JSON data from your web scraping script. Here's the breakdown:
Original Python Script
This is the script you're using to pull data from the Dunelm product page:
import requests import json from bs4 import BeautifulSoup import re from requests_html import HTMLSession url = 'https://www.dunelm.com/product/caldonia-check-natural-eyelet-curtains-1000187301?defaultSkuId=30729125' r = requests.get(url) source_text = r.text # Regex for extract info product_list = re.findall('{\"delivery\"*.*false*}}}', source_text) print(product_list, type((product_list))) with open("json-pattern.json", "w", encoding='utf-8') as file: file.write(str(product_list))
Right now, the product_list variable (a list) has extra characters that are breaking the JSON structure. Let's fix that with the required cleanup rules.
Cleanup Requirements
We need to apply these steps to get valid JSON that can be converted to a Python dictionary:
- Remove the surrounding
['and']entirely - Strip out all instances of
\\"(double backslash + quoted sequence) - Remove all
\\'(backslash + single quote) - Replace every
undefinedwith"undefined" - Ensure there are no spaces between characters after handling steps 3 and 4
Solution Code
Here's the modified script that handles all the cleanup steps properly:
import requests import json import re url = 'https://www.dunelm.com/product/caldonia-check-natural-eyelet-curtains-1000187301?defaultSkuId=30729125' r = requests.get(url) source_text = r.text # Extract the target data product_list = re.findall('{\"delivery\".*false*}}}', source_text) # Check if we found any data before proceeding if product_list: # Step 1: Remove surrounding [' and '] by taking the first element directly cleaned_data = product_list[0] # Step 2: Remove \\" cleaned_data = cleaned_data.replace('\\"', '') # Step 3: Remove \\' cleaned_data = cleaned_data.replace("\\'", "") # Step 4: Replace undefined with "undefined" cleaned_data = cleaned_data.replace('undefined', '"undefined"') # Step 5: Remove any remaining spaces to meet the requirement cleaned_data = cleaned_data.replace(' ', '') # Validate and convert to a Python dictionary try: # Convert HTML quotes to actual JSON quotes first valid_json = cleaned_data.replace('"', '"') product_dict = json.loads(valid_json) print("Successfully converted to dictionary!") print(product_dict) # Save the cleaned, valid JSON to file with open("cleaned-product-data.json", "w", encoding='utf-8') as file: json.dump(product_dict, file, indent=4, ensure_ascii=False) except json.JSONDecodeError as e: print(f"JSON decode error: {e}") print("Cleaned data that caused the error:", cleaned_data) else: print("No matching data found in the page source.")
Key Notes:
- I removed
requests_htmlsince it wasn't being used in the original script (keeps things lean) - Added error handling to catch JSON decode issues, which helps debug if something goes wrong
- The final step converts the cleaned string to a proper Python dictionary, so you can work with the data directly instead of just a raw string
- The saved file uses
json.dump()to write formatted, valid JSON instead of a plain string
内容的提问来源于stack exchange,提问作者MarkWP
相关产品推荐
相关产品推荐

