为何CFS算法的task_tick_fair会在系统tick时调用for_each_sched_entity?
for_each_sched_entity is Called in task_tick_fair for CFS Great question—this cuts right to how Linux's Completely Fair Scheduler (CFS) handles hierarchical scheduling (think cgroup-based task groups). Let's unpack why we can't just update the current task's scheduling entity alone.
The Hierarchy of Scheduling Entities
CFS doesn't only schedule individual tasks—it also manages groups of tasks (via the cgroup subsystem). Every task has its own sched_entity, but each task group it belongs to also has a sched_entity that lives above it in a hierarchy. When a task runs, it's not just using CPU time for itself; it's also consuming the CPU quota allocated to all the groups in its chain.
Why We Need to Traverse the Entire Hierarchy
The task_tick_fair function uses for_each_sched_entity to iterate through every sched_entity in the current task's hierarchy, calling entity_tick on each one. Here's why each level matters:
- Task-level entity: We need to update the task's
vruntime(virtual runtime) to track its CPU usage. This is the core of CFS's fairness logic—tasks with lowervruntimeget priority to run next. - Group-level entities: Each parent group has its own
vruntimeand CPU allocation. When a task in the group runs, we have to increment the group's consumed CPU time too. This ensures groups get their fair share relative to other groups at the same level. For example, if you have two groups each allocated 50% of the CPU, CFS needs to track when one group is using more than its quota to rebalance fairness at the group level.
What If We Skipped This?
If we only updated the task's own sched_entity, group-level fairness would completely break. Groups could exceed their CPU limits, and the guarantees provided by cgroup CPU controllers (like limiting a container's CPU usage) would stop working. The hierarchy exists to enforce fairness at every level—from individual tasks up to entire workload groups.
内容的提问来源于stack exchange,提问作者Lucifer Poltergeist

