无HTML标签时的网页爬取:如何提取spStart函数中的目标数据?
Got it, let's break down how to fix this issue since your target data is tucked away inside the spStart function's parameters instead of regular HTML tags.
Why your previous attempts didn't work
- Beautiful Soup: It only parses HTML elements, so it can't dig into raw JavaScript code blocks where your data lives.
- Selenium: Even though it loads the page, if the
spStartfunction's parameters aren't rendered into visible DOM elements (like divs or spans), thepage_sourcewill just be the original HTML + JS code—no difference from what you get with a simple requests call.
Step-by-step solution
The key is to extract the spStart function's parameters directly from the page source, then parse them to get your altitude and time data. Here's how to do it in Python:
Grab the page source
Use eitherrequests(for static pages) or Selenium (if the page requires JS to load the source itself) to get the full HTML/JS content.Use regex to match the
spStartfunction
Write a regular expression to target the function and pull out its parameters. We'll usere.DOTALLto handle line breaks inside the function.Parse the extracted parameters
Depending on how the parameters are formatted (JSON object, comma-separated values, etc.), parse them to extract the specific data you need.
Example code
import requests import re import json # Replace with your target URL target_url = "https://your-target-site.com" # Get the page source response = requests.get(target_url) page_content = response.text # Regex pattern to find spStart and its parameters # Adjust the pattern if needed based on the actual function syntax pattern = re.compile(r'spStart\((.*?)\);', re.DOTALL) match_result = pattern.search(page_content) if match_result: params_content = match_result.group(1) # Case 1: Parameters are a JSON object (e.g., spStart({"alt": 1500, "open": "08:00", "close": "17:30"});) try: data = json.loads(params_content) altitude = data.get("alt") # Replace with the actual key from your data start_time = data.get("open") end_time = data.get("close") print(f"海拔: {altitude} 米") print(f"通行开始时间: {start_time}") print(f"通行结束时间: {end_time}") except json.JSONDecodeError: # Case 2: Parameters are comma-separated values (e.g., spStart("1500", "08:00", "17:30");) params_list = [param.strip().strip('"\'') for param in params_content.split(',')] if len(params_list) >= 3: altitude = params_list[0] start_time = params_list[1] end_time = params_list[2] print(f"海拔: {altitude} 米") print(f"通行开始时间: {start_time}") print(f"通行结束时间: {end_time}") else: print("参数格式不符合预期,请检查正则表达式和参数结构") else: print("未找到spStart函数,请确认函数名是否正确")
Tips for adjustment
- If the
spStartfunction has more complex formatting (like nested objects or escaped characters), you might need to tweak the regex pattern to ensure you capture all parameters correctly. - If you're using Selenium instead of requests, just replace the
requests.getpart withdriver.page_sourceafter loading the page.
内容的提问来源于stack exchange,提问作者Lorcank11

