如何用BeautifulSoup从HTML的CDATA中提取redshift值
Got it, let's break this down simply. You’ve already nailed the first part by finding the right script tag (index 21)—now we just need to dig out the JSON data hiding in that messy text and pull out the redshift value. Here’s how to do it properly:
Step 1: Grab the script tag’s text content
First, isolate the actual text inside that script tag instead of printing the whole tag object. Update your code like this:
script_content = soup.find_all('script')[21].string
Step 2: Extract and parse the JSON block
That script is probably stuffed with jQuery/JS code, but your target data is in a JSON structure wrapped in curly braces {...}. We’ll use regex to yank out that JSON chunk, then parse it with Python’s built-in json module—way more reliable than hacking strings manually.
Add these lines to your code:
import re import json # Match the full JSON object (handles nested braces with the DOTALL flag) json_pattern = re.compile(r'(\{.*\})', re.DOTALL) match_result = json_pattern.search(script_content) if match_result: # Convert the matched string to a Python dict data_dict = json.loads(match_result.group(1)) # Pull out the redshift value—adjust the path if it's nested in a sub-key redshift_value = data_dict.get('redshift') # If it's nested, like inside an "object" key, use this instead: # redshift_value = data_dict.get('object', {}).get('redshift') print(f"Found redshift: {redshift_value}") else: print("Couldn't locate the JSON data block in the script")
Step 3: Handle CDATA wrappers (if needed)
If the script has CDATA comments like /* <![CDATA[ */ and /* ]]> */ around the JSON, clean those first before matching:
# Strip CDATA comment markers cleaned_script = re.sub(r'/\* <\!\[CDATA\[ \*/|/\* \]\]> \*/', '', script_content) # Then run the regex match on cleaned_script instead
Why this works better than other methods?
- Regex cuts through the noise of extra JS code to target only the JSON you care about
- Using
json.loads()ensures you handle valid JSON correctly, even if the structure has nested keys - It’s flexible—if the redshift is buried in a sub-object, just tweak the
get()path after printingdata_dictto see the full structure
Just print data_dict once to inspect the exact hierarchy, and you’ll be able to grab that 0.06 value in no time.
内容的提问来源于stack exchange,提问作者astrochris

