You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CellDisk故障无法更换,如何启动Oracle Exadata 11gR2 RAC数据库?

Alright, let's walk through your Exadata 11gR2 RAC problem clearly—you're dealing with a faulty SAS CellDisk on an X4720 M3, can't rebuild it right now, and need to get the database up running by skipping this disk. Let's break down your questions one by one:

1. Do you need to delete the PhysicalDisk, and how to use srvctl start database?

First off: you do NOT need to delete the PhysicalDisk. PhysicalDisks map directly to the underlying hardware drives; deleting it would erase critical hardware metadata, making it far harder to rebuild everything once you get a replacement disk. Instead, focus on isolating the faulty CellDisk/Griddisk from ASM and the storage cell layer.

Here's the step-by-step workflow:

Step 1: Clean up the faulty disk in ASM

Log into the ASM instance on one of your database nodes:

sqlplus / as sysasm
  • Identify the faulty disk by matching it to the problematic CellDisk path:
    SELECT name, state, path FROM v$asm_disk WHERE path LIKE '%<your_faulty_celldisk_path>%';
    
  • Mark the disk as offline and drop it immediately (since recovery isn't possible right now):
    ALTER DISKGROUP <your_diskgroup_name> OFFLINE DISK '<asm_disk_name>' DROP AFTER 0;
    
  • Verify the disk is removed and the diskgroup is in a healthy state:
    SELECT name, state FROM v$asm_diskgroup;
    

Step 2: Disable the faulty CellDisk on the storage cell

Log into the affected storage cell and launch cellcli:

cellcli
  • List all failed CellDisks to confirm your target:
    LIST celldisk WHERE status='failed';
    
  • Set the faulty CellDisk to INACTIVE (this preserves hardware records but prevents ASM from attempting to use it):
    ALTER celldisk <faulty_celldisk_name> inactive;
    
  • Confirm the status change took effect:
    LIST celldisk detail WHERE name='<faulty_celldisk_name>';
    

Step 3: Start the database with srvctl

Once ASM is cleaned up, you can start the database normally:

  • First, ensure ASM is running (if it stopped due to the disk issue):
    srvctl start asm -n <database_node_name>
    
  • Start the entire RAC database:
    srvctl start database -d <your_database_name>
    
  • If you only need to start a specific instance (e.g., one node at a time):
    srvctl start instance -d <your_database_name> -i <instance_name>
    

2. How to start the database when there's an active CellDisk failure?

The approach varies based on whether ASM has already recognized the failure and if your diskgroup redundancy is still satisfied. Here's how to handle common scenarios:

Scenario 1: ASM can mount the diskgroup (redundancy holds)

If your diskgroup has enough healthy disks to meet its redundancy requirement (e.g., normal redundancy has at least one mirror copy, high redundancy has two), you can skip straight to starting the database with the srvctl commands above. ASM will automatically ignore the failed disk once it's marked offline/dropped.

Scenario 2: ASM can't mount the diskgroup due to the failed disk

If ASM is stuck trying to access the faulty disk:

  1. Force-mount the ASM instance first:
    sqlplus / as sysasm
    STARTUP FORCE MOUNT;
    
  2. Follow the steps in Section 1 to offline/drop the faulty disk from the diskgroup.
  3. Remount the diskgroup to confirm it's healthy:
    ALTER DISKGROUP <your_diskgroup_name> MOUNT;
    
  4. Now start the database with srvctl as usual.

Scenario 3: Redundancy is insufficient (last copy of data on the faulty disk)

⚠️ Critical Warning: This carries high risk of data loss. Only proceed if you have no other option:

  • Force-mount the diskgroup, ignoring redundancy checks:
    ALTER DISKGROUP <your_diskgroup_name> MOUNT FORCE;
    
  • Once mounted, start the database and immediately back up all critical data. You'll need to replace the faulty disk and restore redundancy as soon as possible.

Key Notes

  • Always back up your ASM metadata (use asmcmd spbackup) and database SPFILE before making changes—this acts as a safety net if something goes wrong.
  • Keep a detailed record of all commands you run; this will simplify rebuilding the CellDisk/Griddisk once you get a replacement disk.
  • Use cellcli list physicaldisk to confirm if the issue is with the CellDisk (logical layer) or PhysicalDisk (hardware layer). If it's a PhysicalDisk failure, the cell will automatically detect a replacement when installed, but leaving it as-is for now is safer than deleting it.

内容的提问来源于stack exchange,提问作者Melih

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:37:28