You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup从HTML的CDATA中提取redshift值

提取目标Redshift值的最优方案

Got it, let's break this down simply. You’ve already nailed the first part by finding the right script tag (index 21)—now we just need to dig out the JSON data hiding in that messy text and pull out the redshift value. Here’s how to do it properly:

Step 1: Grab the script tag’s text content

First, isolate the actual text inside that script tag instead of printing the whole tag object. Update your code like this:

script_content = soup.find_all('script')[21].string

Step 2: Extract and parse the JSON block

That script is probably stuffed with jQuery/JS code, but your target data is in a JSON structure wrapped in curly braces {...}. We’ll use regex to yank out that JSON chunk, then parse it with Python’s built-in json module—way more reliable than hacking strings manually.

Add these lines to your code:

import re
import json

# Match the full JSON object (handles nested braces with the DOTALL flag)
json_pattern = re.compile(r'(\{.*\})', re.DOTALL)
match_result = json_pattern.search(script_content)

if match_result:
    # Convert the matched string to a Python dict
    data_dict = json.loads(match_result.group(1))
    # Pull out the redshift value—adjust the path if it's nested in a sub-key
    redshift_value = data_dict.get('redshift')
    # If it's nested, like inside an "object" key, use this instead:
    # redshift_value = data_dict.get('object', {}).get('redshift')
    
    print(f"Found redshift: {redshift_value}")
else:
    print("Couldn't locate the JSON data block in the script")

Step 3: Handle CDATA wrappers (if needed)

If the script has CDATA comments like /* <![CDATA[ */ and /* ]]> */ around the JSON, clean those first before matching:

# Strip CDATA comment markers
cleaned_script = re.sub(r'/\* <\!\[CDATA\[ \*/|/\* \]\]> \*/', '', script_content)
# Then run the regex match on cleaned_script instead

Why this works better than other methods?

  • Regex cuts through the noise of extra JS code to target only the JSON you care about
  • Using json.loads() ensures you handle valid JSON correctly, even if the structure has nested keys
  • It’s flexible—if the redshift is buried in a sub-object, just tweak the get() path after printing data_dict to see the full structure

Just print data_dict once to inspect the exact hierarchy, and you’ll be able to grab that 0.06 value in no time.

内容的提问来源于stack exchange,提问作者astrochris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 08:11:14