You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python技术问询:如何移除或替换列表中的冗余字符与内容以构建合法字典

Hey there! Let's work through cleaning up that messy JSON data from your web scraping script. Here's the breakdown:

Original Python Script

This is the script you're using to pull data from the Dunelm product page:

import requests
import json
from bs4 import BeautifulSoup
import re
from requests_html import HTMLSession

url = 'https://www.dunelm.com/product/caldonia-check-natural-eyelet-curtains-1000187301?defaultSkuId=30729125'
r = requests.get(url)
source_text = r.text
# Regex for extract info
product_list = re.findall('{\"delivery\"*.*false*}}}', source_text)
print(product_list, type((product_list)))
with open("json-pattern.json", "w", encoding='utf-8') as file:
    file.write(str(product_list))

Right now, the product_list variable (a list) has extra characters that are breaking the JSON structure. Let's fix that with the required cleanup rules.

Cleanup Requirements

We need to apply these steps to get valid JSON that can be converted to a Python dictionary:

  • Remove the surrounding [' and '] entirely
  • Strip out all instances of \\" (double backslash + quoted sequence)
  • Remove all \\' (backslash + single quote)
  • Replace every undefined with "undefined"
  • Ensure there are no spaces between characters after handling steps 3 and 4
Solution Code

Here's the modified script that handles all the cleanup steps properly:

import requests
import json
import re

url = 'https://www.dunelm.com/product/caldonia-check-natural-eyelet-curtains-1000187301?defaultSkuId=30729125'
r = requests.get(url)
source_text = r.text

# Extract the target data
product_list = re.findall('{\"delivery\".*false*}}}', source_text)

# Check if we found any data before proceeding
if product_list:
    # Step 1: Remove surrounding [' and '] by taking the first element directly
    cleaned_data = product_list[0]
    
    # Step 2: Remove \\"
    cleaned_data = cleaned_data.replace('\\"', '')
    
    # Step 3: Remove \\'
    cleaned_data = cleaned_data.replace("\\'", "")
    
    # Step 4: Replace undefined with "undefined"
    cleaned_data = cleaned_data.replace('undefined', '"undefined"')
    
    # Step 5: Remove any remaining spaces to meet the requirement
    cleaned_data = cleaned_data.replace(' ', '')
    
    # Validate and convert to a Python dictionary
    try:
        # Convert HTML quotes to actual JSON quotes first
        valid_json = cleaned_data.replace('"', '"')
        product_dict = json.loads(valid_json)
        print("Successfully converted to dictionary!")
        print(product_dict)
        
        # Save the cleaned, valid JSON to file
        with open("cleaned-product-data.json", "w", encoding='utf-8') as file:
            json.dump(product_dict, file, indent=4, ensure_ascii=False)
    except json.JSONDecodeError as e:
        print(f"JSON decode error: {e}")
        print("Cleaned data that caused the error:", cleaned_data)
else:
    print("No matching data found in the page source.")

Key Notes:

  • I removed requests_html since it wasn't being used in the original script (keeps things lean)
  • Added error handling to catch JSON decode issues, which helps debug if something goes wrong
  • The final step converts the cleaned string to a proper Python dictionary, so you can work with the data directly instead of just a raw string
  • The saved file uses json.dump() to write formatted, valid JSON instead of a plain string

内容的提问来源于stack exchange,提问作者MarkWP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 07:04:06