使用Selenium Python爬取巴基斯坦论坛报新闻的技术问题求助
Selenium Python爬取巴基斯坦论坛新闻的问题修复
问题1:提取新闻链接时重复获取href
你的代码中//div[contains(@class,'flex-wrap')]//a会匹配该容器下所有<a>标签,而每个新闻卡片通常包含图片和标题两个<a>(均指向同一新闻链接),导致重复提取。
修复方案
调整XPATH定位,只抓取每个新闻卡片内的第一个<a>,或者直接定位标题对应的<a>(更准确);也可通过集合去重避免重复链接:
def Url_Extraction(): category_name = driver.find_element(By.XPATH, '//*[@id="main-section"]/h1') cat = category_name.text print(f"{cat}") # 修改XPATH,定位每个新闻卡片的第一个<a> news_articles = driver.find_elements(By.XPATH,"//div[contains(@class,'flex-wrap')]/div//a[1]") # 用集合去重,避免重复链接 seen_urls = set() for element in news_articles: URL = element.get_attribute('href') if URL not in seen_urls: seen_urls.add(URL) Url.append(URL) Category.append(cat) current_time = time.time() - start_time print(f'{len(Url)} urls extracted') print(f'{len(Category)} categories extracted') print(f'Current Time: {current_time / 3600:.2f} hr, {current_time / 60:.2f} min, {current_time:.2f} sec', flush=True)
问题2:点击新闻链接无法获取完整文章内容
常见原因及修复方法:
- 新窗口打开未切换:点击链接后若打开新窗口,需切换到新窗口再提取内容
- 动态内容未等待加载:文章内容多为异步加载,需显式等待元素出现
- 滚动加载内容:部分长文需要滚动页面才能加载完整内容
修复示例(点击链接后提取内容)
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time def extract_article_content(url): original_window = driver.current_window_handle # 打开目标链接 driver.get(url) # 等待文章内容容器加载完成 WebDriverWait(driver, 15).until( EC.presence_of_element_located((By.XPATH, "//div[contains(@class, 'story-detail')]")) ) # 滚动到底部触发剩余内容加载 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) # 提取完整文章内容 content = driver.find_element(By.XPATH, "//div[contains(@class, 'story-detail')]").text return content
内容的提问来源于stack exchange,提问作者SanaSajid
相关产品推荐
相关产品推荐

