UBIFS文件系统创建后读取错误问题及相关技术疑问
Problem Overview
When creating a new UBIFS filesystem on an existing UBI partition's volume and running tar -cf /dev/null * to read all files, we hit a critical UBIFS error. The kernel error log is as follows:
UBIFS error (pid 1810): ubifs_read_node: bad node type (255 but expected 1) UBIFS error (pid 1810): ubifs_read_node: bad node at LEB 33:967120, LEB mapping status 0 Not a node, first 24 bytes: 00000000: ff ff ff ff ff ff ff ff ff ff ff ff ff ff ff ff ff ff ff ff ff ff ff ff CPU: 0 PID: 1810 Comm: tar Tainted: P O 3.18.80 #3 Backtrace: [<8001b664>] (dump_backtrace) from [<8001b880>] (show_stack+0x18/0x1c) r7:00000000 r6:00000000 r5:80000013 r4:80566cbc [<8001b868>] (show_stack) from [<801d77e0>] (dump_stack+0x88/0xa4) [<801d7758>] (dump_stack) from [<80171608>] (ubifs_read_node+0x1fc/0x2b8) r7:00000021 r6:00000712 r5:000ec1d0 r4:bd59d000 [<8017140c>] (ubifs_read_node) from [<8018b654>](ubifs_tnc_read_node+0x88/0x124) r10:00000000 r9:b4f9fdb0 r8:bd59d264 r7:00000001 r6:b71e6000 r5:bd59d000 r4:b6126148 [<8018b5cc>] (ubifs_tnc_read_node) from [<80174690>](ubifs_tnc_locate+0x108/0x1e0) r7:00000001 r6:b71e6000 r5:00000001 r4:bd59d000 [<80174588>] (ubifs_tnc_locate) from [<80168200>] (do_readpage+0x1c0/0x39c) r10:bd59d000 r9:000005b1 r8:00000000 r7:bd258710 r6:a5914000 r5:b71e6000 r4:be09e280 [<80168040>] (do_readpage) from [<80169398>] (ubifs_readpage+0x44/0x424) r10:00000000 r9:00000000 r8:bd59d000 r7:be09e280 r6:00000000 r5:bd258710 r4:00000000 [<80169354>] (ubifs_readpage) from [<80089654>](generic_file_read_iter+0x48c/0x5d8) r10:00000000 r9:00000000 r8:00000000 r7:be09e280 r6:bd2587d4 r5:bd5379e0 r4:00000000 [<800891c8>] (generic_file_read_iter) from [<800bbd60>](new_sync_read+0x84/0xa8) r10:00000000 r9:b4f9e000 r8:80008e24 r7:bd6437a0 r6:bd5379e0 r5:b4f9ff80 r4:00001000 [<800bbcdc>] (new_sync_read) from [<800bc85c>] (__vfs_read+0x20/0x54) r7:b4f9ff80 r6:7ec2d800 r5:00001000 r4:800bbcdc [<800bc83c>] (__vfs_read) from [<800bc91c>] (vfs_read+0x8c/0xf4) r5:00001000 r4:bd5379e0 [<800bc890>] (vfs_read) from [<800bc9cc>] (SyS_read+0x48/0x80) r9:b4f9e000 r8:80008e24 r7:00001000 r6:7ec2d800 r5:bd5379e0 r4:bd5379e0 [<800bc984>] (SyS_read) from [<80008c80>] (ret_fast_syscall+0x0/0x3c) r7:00000003 r6:7ec2d800 r5:00000008 r4:00082a08 UBIFS error (pid 1811): do_readpage: cannot read page 0 of inode 1457, error -22
The mount and UBI attachment details:
UBI-0: ubi_attach_mtd_dev:default fastmap pool size: 190 UBI-0: ubi_attach_mtd_dev:default fastmap WL pool size: 25 UBI-0: ubi_attach_mtd_dev:attaching mtd3 to ubi0 UBI-0: scan_all:scanning is finished UBI-0: ubi_attach_mtd_dev:attached mtd3 (name "data", size 3824 MiB) UBI-0: ubi_attach_mtd_dev:PEB size: 1048576 bytes (1024 KiB), LEB size: 1040384 bytes UBI-0: ubi_attach_mtd_dev:min./max. I/O unit sizes: 4096/4096, sub-page size 4096 UBI-0: ubi_attach_mtd_dev:VID header offset: 4096 (aligned 4096), data offset: 8192 UBI-0: ubi_attach_mtd_dev:good PEBs: 3816, bad PEBs: 8, corrupted PEBs: 0 UBI-0: ubi_attach_mtd_dev:user volume: 6, internal volumes: 1, max. volumes count: 128 UBI-0: ubi_attach_mtd_dev:max/mean erase counter: 4/2, WL threshold: 4096, image sequence number: 1362948729 UBI-0: ubi_attach_mtd_dev:available PEBs: 2313, total reserved PEBs: 1503, PEBs reserved for bad PEB handling: 72 UBI-0: ubi_thread:background thread "ubi_bgt0d" started, PID 419 UBIFS: mounted UBI device 0, volume 6, name "slot1", R/O mode UBIFS: LEB size: 1040384 bytes (1016 KiB), min./max. I/O unit sizes: 4096 bytes/4096 bytes UBIFS: FS size: 259055616 bytes (247 MiB, 249 LEBs), journal size 12484608 bytes (11 MiB, 12 LEBs) UBIFS: reserved for root: 4952683 bytes (4836 KiB) UBIFS: media format: w4/r0 (latest is w4/r0), UUID 92C7B251-2666-4717-B735-5539900FE749, small LPT model
Reproduction Context
- Reproduction rate: ~1 in 20 operation sequences
- Operation sequence:
ubirmvolubimkvol(create 256MB volume)- Unpack
rootfs.tgz(20MB compressed, 40MB unpacked) onto mounted UBIFS umountubidetachsyncreboot
- Environment: ARM board running Linux 3.18.80, issue occurs across boards with different vendor NAND chips
- Attempted fixes:
- Disabled
CONFIG_MTD_UBI_FASTMAP(no improvement) - Applied the "ubifs: Fix journal replay w.r.t. xattr nodes" patch (no improvement)
- Disabled
Questions & Detailed Answers
1. Are there special operations required before reboot besides umount, ubidetach, and sync? Or any free space repair requirements?
Yes, there are a few additional steps you can take to ensure the UBIFS/UBI stack is in a consistent state before reboot:
- Remount as read-only first: Before running
umount, usemount -o remount,ro /path/to/ubifs. This forces UBIFS to flush all dirty data, commit the journal, and mark the filesystem as clean. It's more reliable than a regular umount for ensuring no pending writes are left. - Clear page cache: Run
echo 3 > /proc/sys/vm/drop_cachesto free page cache entries, then runsyncagain. This ensures any cached data is written to the storage device. - Wait for UBI background threads: The
ubi_bgt0dthread handles wear-leveling and cleanup tasks. You can wait for it to finish its work by checking thestatefile in/sys/class/ubi/ubi0/(wait until it showsidle). Some systems include anubi-watchtool to automate this wait. - Verify UBI detachment: After
ubidetach, check that/dev/ubi0and related volume nodes are removed. If they persist, you may need to wait a few seconds or trigger a rescan before rebooting.
UBIFS doesn't require explicit free space repair in normal operation, but ensuring all write operations are fully committed before reboot eliminates potential consistency issues.
2. Will creating the filesystem image offline with mkfs.ubifs instead of unpacking via tar fix the issue?
This is a strong candidate for resolving the problem, and it's worth testing immediately:
- Why it helps: When you unpack
taronto a mounted UBIFS, the filesystem performs real-time allocation, metadata updates, and journal writes. In Linux 3.18's UBIFS implementation, there may be edge cases (like rapid sequential writes or specific metadata handling) that lead to inconsistent node states after reboot. - Offline workflow:
- Create a temporary directory and unpack
rootfs.tgzinto it. - Use
mkfs.ubifsto create an offline image, matching your UBI volume's LEB size and other parameters:mkfs.ubifs -d /tmp/rootfs -e 1040384 -c 249 -o rootfs.img - Attach the UBI device, create the volume, then write the image with
ubiupdatevol:ubiattach /dev/ubi_ctrl -m 3 ubimkvol /dev/ubi0 -N slot1 -s 256MiB ubiupdatevol /dev/ubi0_6 rootfs.img
- Create a temporary directory and unpack
- Benefit: Offline creation builds a consistent filesystem structure upfront, avoiding the incremental write path that may be triggering the bug. If the error disappears with this method, it confirms the issue is tied to online filesystem modification.
3. Will backporting fs/ubifs commits from Linux v4.14 help resolve the issue?
Absolutely—backporting relevant UBIFS fixes from Linux 4.14 is a high-impact troubleshooting step:
- Context: Linux 3.18 is an older LTS release, and the UBIFS subsystem received dozens of stability fixes between 3.18 and 4.14, including fixes for:
- Journal replay inconsistencies
- Node type validation errors (exactly the error you're seeing)
- Fastmap-related corruption
- LEB allocation edge cases
- Best practices:
- Focus on commits related to
ubifs_read_node, journal replay, metadata validation, and fastmap. Avoid cherry-picking single commits; instead, port a coherent set of fixes to avoid dependency issues. - Test incrementally: Port a small batch of fixes, rebuild the kernel, and check if the reproduction rate drops or disappears.
- Focus on commits related to
- Caveat: Some commits may depend on other kernel changes (e.g., VFS or MTD subsystem updates), so you'll need to adapt patches to your 3.18 tree. However, many UBIFS-specific fixes are self-contained and can be ported with minimal effort.
内容的提问来源于stack exchange,提问作者patraulea

