网页爬取问题:使用.find方法无法找到页面中的指定字符串
网页爬取问题:无法在script标签内定位目标字符串
我正在编写第一个Python程序,做网页爬取时遇到了问题。目标字符串thisstring在页面的script标签里,结构如下:
<script> anotherstring; thisstring = {...}; </script>
我的代码如下:
import requests from bs4 import BeautifulSoup page = requests.get('www.somewebadress.com') soup = BeautifulSoup(page.content, 'html.parser') lines = soup.find_all('script') x = 0 # 用来统计html里script标签的数量,计数是对的 for line in lines: x = x + 1 txt = line.find('thisstring') # 用"thisstring"也没用 if txt == None: print("not found") else: print("found") print(x)
明明用print(line)能看到目标字符串,但怎么都匹配不到,试了各种网上的方法都没用,折腾了一整天。用的是Spyder,会不会是这个的问题?
解决方法
核心问题:错误使用find()方法
BeautifulSoup的find()是用来查找HTML子标签的,不是搜索文本内容。要检查script标签里的文本是否包含目标字符串,直接用字符串包含判断就行:'thisstring' in line.text。
次要问题:URL缺少协议头
requests.get('www.somewebadress.com')会请求失败,必须加上http://或https://,比如https://www.somewebadress.com。
修正后的代码
import requests from bs4 import BeautifulSoup # 补充协议头 page = requests.get('https://www.somewebadress.com') soup = BeautifulSoup(page.content, 'html.parser') scripts = soup.find_all('script') counter = 0 for script in scripts: counter += 1 if 'thisstring' in script.text: print(f"在第{counter}个script标签中找到目标字符串") else: print(f"第{counter}个script标签中未找到") print(f"共找到{counter}个script标签")
进阶:提取thisstring对应的JSON数据
如果需要把thisstring = {...}里的{...}内容提取出来,可以用正则表达式匹配:
import requests from bs4 import BeautifulSoup import re import json page = requests.get('https://www.somewebadress.com') soup = BeautifulSoup(page.content, 'html.parser') scripts = soup.find_all('script') for script in scripts: if 'thisstring' in script.text: # 匹配thisstring = 后面的JSON块 match_result = re.search(r'thisstring = ({.*?});', script.text, re.DOTALL) if match_result: json_data = json.loads(match_result.group(1)) print("提取到的数据:", json_data)
关于Spyder
Spyder不是问题根源,工具本身没问题,只是代码逻辑需要调整。
内容的提问来源于stack exchange,提问作者My Ka
相关产品推荐
相关产品推荐

