You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无法抓取工单网站数据,求Beautiful Soup爬虫问题排查方案

问题分析与解决方案

核心问题排查

先确定页面数据是静态还是动态:

  • 执行以下代码,把请求到的页面保存为HTML,打开查看是否包含目标表格数据:
import requests
url = "https://newtin.co811.org/responsedisplay/?ticket=B420800874"
response = requests.get(url)
with open("test_page.html", "w", encoding="utf-8") as f:
    f.write(response.text)
  • 如果打开test_page.html能看到目标数据:说明是CSS选择器错误,需要修正选择器。
  • 如果看不到目标数据:说明是动态数据加载,页面内容由JavaScript渲染,requests无法直接获取,需要用浏览器模拟工具。

解决方案1:静态数据场景(修正CSS选择器)

  1. 打开目标页面,按F12打开开发者工具。
  2. 定位到要提取的表格元素,右键选择「复制」→「复制CSS选择器」。
  3. 替换代码中get_ticket_data函数里的占位选择器:
# 示例:假设目标表格的CSS选择器是table.ticket-details
table = soup.select_one("table.ticket-details")
if table:
    # 提取表格内的具体数据,遍历行和列
    rows = table.find_all("tr")
    table_data = []
    for row in rows:
        cols = row.find_all("td")
        cols_text = [col.text.strip() for col in cols]
        table_data.append(cols_text)
else:
    table_data = "N/A"
  1. 重新运行代码,即可提取到正确的表格数据。

解决方案2:动态数据场景(使用Selenium模拟浏览器)

步骤1:安装依赖

pip install selenium webdriver-manager

步骤2:修改get_ticket_data函数

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.chrome.options import Options
import time

def get_ticket_data(ticket_number):
    ticket_without_last_4 = ticket_number[:-4]
    url = f"https://newtin.co811.org/responsedisplay/?ticket={ticket_without_last_4}"
    
    # 配置Chrome无头模式(不显示浏览器窗口)
    chrome_options = Options()
    chrome_options.add_argument("--headless=new")
    chrome_options.add_argument("--disable-gpu")
    
    # 初始化浏览器
    driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=chrome_options)
    driver.get(url)
    # 等待页面加载完成(根据实际情况调整等待时间)
    time.sleep(2)
    
    # 获取渲染后的页面源码
    soup = BeautifulSoup(driver.page_source, 'html.parser')
    driver.quit()
    
    # 提取表格数据(替换为实际的CSS选择器)
    table = soup.select_one("table.ticket-details")
    table_data = []
    if table:
        rows = table.find_all("tr")
        for row in rows:
            cols = row.find_all(["td", "th"])
            cols_text = [col.text.strip() for col in cols]
            table_data.append(cols_text)
    else:
        table_data = "N/A"
    
    return {"Table Data": table_data}

步骤3:调整process_excel_file函数中的数据处理逻辑

根据返回的表格数据格式,修改data_list.append中的字段:

data_list.append({
    "Ticket": ticket,
    "Account": row['Account'],
    "Completed": row['Completed'],
    "Work Date": row['Work Date'],
    "Type": row['Type'],
    "Category": row['Category'],
    "Priority": row['Priority'],
    "Company Name": row['Company Name'],
    "Web Table Data": str(web_data['Table Data'])  # 或整理成更易读的格式
})

注意事项

  • 频繁请求可能触发网站反爬机制,建议在请求之间添加延迟(比如time.sleep(1))。
  • 确保Excel中的「Ticket」列字段名与代码中的row['Ticket']一致,否则会报错。
  • 如果使用Selenium,webdriver-manager会自动处理Chrome浏览器与ChromeDriver的版本兼容问题。

内容的提问来源于stack exchange,提问作者Robert Meizo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 10:17:33