求助:运行GeeksforGeeks网页爬取词频Python代码无结果
问题排查与修复方案
核心问题1:页面结构变更,目标元素无法匹配
原代码依赖的div.entry-content选择器已不适用于当前目标页面的结构,导致soup.findAll('div', {'class': 'entry-content'})返回空列表,循环体完全未执行,后续词频统计逻辑根本没触发,因此无任何输出。
核心问题2:缺少请求头,可能被反爬拦截
直接用requests.get(url)发起请求,未携带浏览器标识等请求头,部分网站会返回异常内容(如反爬页面、空白内容),同样会导致BeautifulSoup无法解析到目标文本。
修复后的完整代码
# Python3 网页爬取与词频统计程序 import requests from bs4 import BeautifulSoup from collections import Counter def start(url): wordlist = [] # 添加请求头模拟浏览器访问 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36' } try: source_code = requests.get(url, headers=headers).text except Exception as e: print(f"请求出错: {e}") return soup = BeautifulSoup(source_code, 'html.parser') # 适配当前页面结构,选择文章主体内容的容器 for each_text in soup.select('div.article-content'): content = each_text.text words = content.lower().split() wordlist.extend(words) # 移出循环,避免重复处理 clean_wordlist(wordlist) def clean_wordlist(wordlist): clean_list = [] symbols = "!@#$%^&*()_-+={[}]|\;:\"<>?/., " for word in wordlist: for sym in symbols: word = word.replace(sym, '') if len(word) > 0: clean_list.append(word) create_dictionary(clean_list) def create_dictionary(clean_list): word_count = Counter(clean_list) top = word_count.most_common(10) print(top) if __name__ == '__main__': url = "https://www.geeksforgeeks.org/programming-language-choose/" start(url)
关键修复点说明
- 更新页面选择器:将原有的
div.entry-content替换为当前页面实际承载文章内容的div.article-content(可通过浏览器开发者工具查看页面结构确认)。 - 添加请求头:通过
headers参数模拟浏览器请求,避免被网站反爬机制拦截。 - 优化循环逻辑:将
clean_wordlist(wordlist)移出for each_text循环,避免重复处理已收集的词列表,提升效率。 - 简化词频统计:直接用
Counter(clean_list)替代手动构建字典的逻辑,代码更简洁高效。
内容的提问来源于stack exchange,提问作者EdG
相关产品推荐
相关产品推荐

