无法抓取.aspx网站新闻链接的Python爬虫问题求助
.aspx网站新闻链接爬取问题解决
我用Python爬取企业网站新闻稿时,遇到.aspx网站的爬取障碍:换用lxml解析器替代html.parser后,仍只能抓取到联系、招聘这类基础链接,无法获取实际新闻文章链接。观察网页元素,目标新闻链接格式如下:
<a class="module_headline-link" href="/news-and-events/news/news-details/2022/Compugen-to-Release-Second-Quarter-Results-on-Thursday-August-4-2022/default.aspx">Compugen to Release Second Quarter Results on Thursday, August 4, 2022</a>
普通HTML网站用注释内的过滤代码可正常筛选链接,但.aspx网站无法生效,推测是对lxml细节不熟悉,或解析器识别相对路径时出现问题。以下是三个存在问题的.aspx网站爬虫代码,以及一个正常运行的HTML网站代码:
问题案例:三家.aspx网站爬虫代码
COMPANY 1
from bs4 import BeautifulSoup as soup import requests import pandas as pd URL = 'https://ir.cgen.com/news-and-events/news/default.aspx' full = '' html_text = requests.get(URL).text chickennoodle = soup(html_text, 'lxml') for link in chickennoodle.find_all('a'): my_links = (link.get('href')) print(my_links) #if str(my_links).startswith("/news-and-events/news/news-details/"): # print(str(full)+my_links) #else: # None
COMPANY 2
from bs4 import BeautifulSoup as soup import requests import pandas as pd URL = 'https://www.meipharma.com/media/press-releases' full = '' html_text = requests.get(URL).text chickennoodle = soup(html_text, 'html.parser') for link in chickennoodle.find_all('a'): my_links = (link.get('href')) print(my_links) # if str(my_links).startswith(""): # print(str(full)+my_links) # else: # None
COMPANY 3
from bs4 import BeautifulSoup as soup import requests import pandas as pd URL = 'https://investor.sierraoncology.com/news-releases/default.aspx' full = '' html_text = requests.get(URL).text chickennoodle = soup(html_text, 'lxml') for link in chickennoodle.find_all('a'): my_links = (link.get('href')) print(my_links)
正常运行的HTML站点代码
from bs4 import BeautifulSoup as soup import requests import pandas as pd URL = "https://investors.aileronrx.com/index.php/news-releases" full = "https://investors.aileronrx.com" ALRNlinks = [] html_text = requests.get(URL).text chickennoodle = soup(html_text, 'html.parser') for link in chickennoodle.find_all('a'): my_links = (link.get('href')) if str(my_links).startswith("/news-rele"): ALRN = (str(full)+my_links) ALRNlinks.append(ALRN) print(ALRNlinks)
解决方案及修正代码
核心问题分析
lxml解析器完全支持识别以/开头的相对路径,问题出在未精准定位新闻链接的特征(比如类名、href规则),以及缺少基础URL拼接的环节。
COMPANY 1 修正代码
from bs4 import BeautifulSoup as soup import requests import pandas as pd URL = 'https://ir.cgen.com/news-and-events/news/default.aspx' base_url = 'https://ir.cgen.com' # 补充网站基础URL html_text = requests.get(URL).text chickennoodle = soup(html_text, 'lxml') news_links = [] # 通过类名+href规则精准筛选新闻链接 for link in chickennoodle.find_all('a', class_='module_headline-link', href=True): href = link['href'] if href.startswith('/news-and-events/news/news-details/'): full_link = base_url + href news_links.append(full_link) print(full_link) # 简洁写法(CSS选择器) # news_links = [base_url + a['href'] for a in chickennoodle.select('a.module_headline-link[href^="/news-and-events/news/news-details/"]')] # print(news_links)
COMPANY 2 修正代码
from bs4 import BeautifulSoup as soup import requests import pandas as pd URL = 'https://www.meipharma.com/media/press-releases' base_url = 'https://www.meipharma.com' html_text = requests.get(URL).text chickennoodle = soup(html_text, 'lxml') # 统一使用lxml提升解析效率 news_links = [] # 根据页面实际特征筛选:新闻链接href以/media/press-releases/开头 for link in chickennoodle.find_all('a', href=True): href = link['href'] if href.startswith('/media/press-releases/') and 'page=' not in href: # 排除分页链接 full_link = base_url + href news_links.append(full_link) print(full_link)
COMPANY 3 修正代码
from bs4 import BeautifulSoup as soup import requests import pandas as pd URL = 'https://investor.sierraoncology.com/news-releases/default.aspx' base_url = 'https://investor.sierraoncology.com' html_text = requests.get(URL).text chickennoodle = soup(html_text, 'lxml') news_links = [] # 该网站新闻链接href以/news-releases/news-release-details/开头 for link in chickennoodle.find_all('a', href=True): href = link['href'] if href.startswith('/news-releases/news-release-details/'): full_link = base_url + href news_links.append(full_link) print(full_link) # 简洁写法(CSS选择器) # news_links = [base_url + a['href'] for a in chickennoodle.select('a[href^="/news-releases/news-release-details/"]')] # print(news_links)
关键注意事项
- 精准筛选元素:避免遍历所有
<a>标签,通过类名、href属性前缀定位目标链接,减少无效数据。 - 拼接完整URL:以
/开头的相对路径必须和网站基础URL拼接,才能得到可访问的完整链接。 - 空值处理:使用
href=True过滤无href属性的标签,避免处理None值报错。 - 解析器统一:优先使用
lxml解析器,其容错性和解析速度优于html.parser。
内容的提问来源于stack exchange,提问作者Ash Cash
相关产品推荐
相关产品推荐

