Linux sysfs文件修改时机与函数及块层属性值变更问询
Hey there, let's break down your questions about Linux block layer sysfs, specifically around blk-mq-sysfs.c, attribute modification mechanisms, and related timing/functions:
First, it's key to understand that sysfs files aren't regular disk files—they're kernel-user space interfaces backed by kernel functions. Every sysfs attribute you see is defined via a struct kobj_attribute, which includes two critical function pointers:
show(): Called when user space reads the attribute (e.g.,catthe file) to fetch the current kernel state.store(): Called when user space writes to the attribute (e.g.,echoa value) to modify the corresponding kernel state.
The sysfs_create_group() function you mentioned is just the way to register a collection of these attributes to a specific kobject (the core structure representing a sysfs node, like a block device's node).
blk-mq-sysfs.c Looking at the blk-mq-sysfs.c file, all the MQ-related sysfs attributes are grouped into blk_mq_attr_group. For example, take the nr_hw_queues attribute (controls the number of hardware queues):
static ssize_t nr_hw_queues_store(struct device *dev, struct device_attribute *attr, const char *buf, size_t count) { struct request_queue *q = dev_to_queue(dev); unsigned int nr; int ret; ret = kstrtouint(buf, 10, &nr); if (ret) return ret; // Core logic to modify the queue's hardware queue count ret = blk_mq_update_nr_hw_queues(q, nr); return ret ? ret : count; } static DEVICE_ATTR_RW(nr_hw_queues);
This nr_hw_queues_store function is the actual entry point where kernel state gets modified when you write to the corresponding sysfs file. The DEVICE_ATTR_RW macro wraps the attribute definition, linking it to both its show and store functions.
There are two main scenarios for sysfs attribute changes:
- User-initiated modification: When you run a command like
echo 4 > /sys/block/nvme0n1/queue/nr_hw_queues, the kernel triggers the attribute'sstorefunction (likenr_hw_queues_storeabove). This function parses the input value, validates it, and calls internal kernel functions (e.g.,blk_mq_update_nr_hw_queues) to adjust the block layer configuration. - Kernel-initiated state updates: For attributes that reflect dynamic kernel state (like I/O statistics), the kernel doesn't "write" to the sysfs file directly. Instead, it updates its internal data structures, then notifies user space of the change using functions like:
sysfs_notify(struct kobject *kobj, const char *dir, const char *attr): Triggers a notification that the attribute's value has changed.kobject_uevent(struct kobject *kobj, enum kobject_action action): Sends a uevent to user space, which tools likeudevcan act on.
When user space reads the attribute after this notification, the show function fetches the latest updated value from the kernel.
For your NVMe 750 device:
- The sysfs attributes are created during the device probe process: when the NVMe driver initializes the request queue, it calls
blk_mq_sysfs_init()(fromblk-mq-sysfs.c), which registers theblk_mq_attr_groupto the queue's device kobject. - When I/O requests are processed, attributes related to I/O stats (like
io_ticksorin_flight) are updated internally by the block layer. User space reads these via their respectiveshowfunctions, which pull the latest stats from kernel structures.
内容的提问来源于stack exchange,提问作者Doodu

