使用Beautiful Soup提取Pine Script参考手册数据遇空JSON问题求助
Pine Script参考手册数据提取:JSON文件为空的解决方法
你尝试用Beautiful Soup提取TradingView Pine Script v5参考手册的数据到JSON文件,但生成的文件始终为空。问题出在两个关键地方:请求头缺失导致页面内容未正确获取,以及CSS类选择器写法错误。
问题分析
- 请求被拦截:TradingView会校验请求的User-Agent,默认
requests库的UA会被识别为爬虫,返回的页面不包含目标内容。 - 类选择器语法错误:
tv-pine-reference-item__text.tv-text是CSS选择器的写法,但BeautifulSoup的find方法中class_参数不能直接用点连接多个类名,需用空格分隔或改用select_one方法。
修改后的代码
import requests from bs4 import BeautifulSoup import json url = "https://www.tradingview.com/pine-script-reference/v5/" # 模拟浏览器请求头,避免被拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } response = requests.get(url, headers=headers) # 校验请求是否成功,失败则抛出异常 response.raise_for_status() soup = BeautifulSoup(response.text, "html.parser") def extract_data(soup): data = {} # 遍历所有参考项容器 for item in soup.find_all("div", class_="tv-pine-reference-item"): h3 = item.find("h3") if not h3: continue topic = h3.text.strip() # 修正类选择器:用空格分隔多个类名 text = item.find("div", class_="tv-pine-reference-item__text tv-text") definition = text.text.strip() if text else None syntax = item.find("pre", class_="tv-pine-reference-item__syntax") example = syntax.text.strip() if syntax else None data[topic] = {"definition": definition, "example": example} return data data = extract_data(soup) # 写入JSON文件,指定编码避免中文乱码,格式化输出便于阅读 with open("PineScriptv5Manual.json", "w", encoding="utf-8") as f: json.dump(data, f, ensure_ascii=False, indent=2)
关键修改说明
- 添加User-Agent:模拟正常浏览器请求,确保获取完整的页面内容。
- 修正类选择器:将多类名的点连接改为空格分隔,适配
find方法的参数规则;也可改用item.select_one("div.tv-pine-reference-item__text.tv-text")直接使用CSS选择器语法。 - 请求状态校验:
response.raise_for_status()在请求失败时抛出异常,快速定位问题。 - 文本优化:用
strip()去除文本前后空白,让JSON数据更整洁。 - JSON输出优化:指定
encoding="utf-8"避免中文乱码,indent=2让JSON结构更清晰。
内容的提问来源于stack exchange,提问作者AinSophAur
相关产品推荐
相关产品推荐

