▶ Cinematic fable · Watch on YouTube ▶ 影片版寓言 · 在 YouTube 观看
城东有一家很忙的饭馆,厨房在后院,大堂在前面,中间隔着一道厚墙,只开了一扇小门。
最早的规矩是这样的:客人点了菜,跑堂的就推开小门进厨房,把菜单递给大厨,站在一旁等。大厨炒好一盘,跑堂的端着走回大堂,放到桌上,再回去等下一盘。一道门来回跑,客人多的时候跑堂的全堵在门口排队——大堂这边喊催菜,厨房那边喊没人来端。
后来老板想了个办法。他把小门封了,在墙上开了一个长方形的传菜口,装上两条首尾相连的环形转盘——一条从大堂转进厨房,一条从厨房转出大堂。
从此规矩变了。
跑堂的不用推门了。他把写好的菜单纸条放在大堂这侧转盘的空格里,转盘自己慢慢转进厨房。厨房那边,切菜的小工盯着转盘,看到纸条就摘下来递给大厨。菜炒好了,小工把盘子放在回程转盘的空格里,转出去。大堂这边,跑堂的看到盘子出来,端走上桌。
没有人走过那道门。
跑堂的动作极快:写纸条、往格子里一塞,转身就去招呼下一桌——他甚至可以一口气塞五六张纸条再回头看回程转盘有没有出菜。厨房那边也一样,大厨按自己的节奏批量取单、批量出菜,不用等跑堂的一趟趟来。
转盘是环形的,格子转完一圈就回到起点。满了就是满了——跑堂的塞不下就等一格空出来再塞。同样,回程转盘满了,厨房也得等跑堂的把盘子端走腾出空格。但只要双方各干各的、不停手,流水从来不断。
聪明的跑堂还发现一件事:他可以提前把下一桌客人可能点的招牌菜纸条塞进去——客人还没开口,转盘里已经排着单了。这样菜一出就能上,客人觉得快得不可思议。
有个新来的跑堂问老伙计:”以前推门进厨房,我一次就只能带一张单,还得站在门口等大厨接。现在我站在大堂就能连塞十张单,这中间省的到底是什么?”
老伙计笑了:”省的是过门。每推一次门,门闩要开、要关、要落锁、要验人——这套手续比炒菜还慢。现在纸条往格子一放,你连门在哪都不用知道。”
——到这儿你大概已经认出来了:这就是 io_uring。
这是什么
io_uring 是 Linux 5.1(2019)引入的异步 I/O 框架。它在用户空间和内核之间共享两个环形缓冲区(ring buffer):提交队列(Submission Queue, SQ)和完成队列(Completion Queue, CQ)。应用把 I/O 请求(读、写、accept、send、recv……)填进 SQ 的空槽,内核从 SQ 取出请求执行,完成后把结果填进 CQ。应用再从 CQ 取结果。
关键优势:无需每次 I/O 都做 syscall。传统 read/write 每调用一次就要 user→kernel→user 切换一次(context switch + 安全检查),在高并发场景下这个开销(”过门”)比实际 I/O 还大。io_uring 把多个请求批量提交(一次 io_uring_enter 甚至 polling 模式下零 syscall),内核批量完成,双方只通过共享内存通信。
为什么重要
在网络密集型(Nginx、数据库 proxy、存储引擎)和磁盘密集型(RocksDB、Ceph)场景中,io_uring 比传统 epoll + 线程池模型快 30%–300%,因为它砍掉了最大的固定开销:syscall 上下文切换。它也是 Linux 上唯一能做到真正异步磁盘 I/O 的通用接口(老的 aio 只支持 direct I/O 且限制多)。如果你运行高吞吐 Kubernetes workload(数据库、消息队列、对象存储),底层引擎大概率已经在用 io_uring。理解它的 ring 语义、backpressure 机制(SQ 满则阻塞)和 polling 模式选择,直接影响你调 nr_requests、iodepth 和 CPU 亲和性时能不能做对。
隐喻对应表
- 厨房在后院,大堂在前面,中间隔一道厚墙 → 内核空间与用户空间的隔离
- 推门进厨房递单、站着等 → 传统
read/writesyscall,一次调用一次上下文切换 - 门闩开关落锁验人 → syscall 的 context switch + 安全检查开销
- 封掉小门,开传菜口装两条转盘 → 用
io_uring_setup创建共享 ring buffer - 大堂到厨房的转盘 → Submission Queue(SQ),应用放入 I/O 请求
- 厨房到大堂的转盘 → Completion Queue(CQ),内核放入完成结果
- 纸条塞进格子 → 往 SQ 填一个 SQE(Submission Queue Entry)
- 盘子放进回程格子 → 内核往 CQ 填一个 CQE(Completion Queue Entry)
- 一口气塞五六张单再回头看 → 批量提交,一次
io_uring_enter提交多个请求 - 转盘环形,满了就等 → ring buffer 固定大小,SQ 满则 backpressure
- 提前塞招牌菜纸条 → 预注册 buffer / fixed file,减少每次请求的设置开销
- 没有人走过那道门 → polling 模式下零 syscall,纯共享内存通信
- 跑堂和大厨各干各的不停手 → 用户空间与内核异步并行,流水线不阻塞
There is a busy restaurant on the east side of town. The kitchen is in the back courtyard, the dining hall up front, and between them stands a thick wall with a single small door.
The old system worked like this: a customer orders, the waiter pushes open the door, walks into the kitchen, hands the slip to the chef, and stands there waiting. When the dish is done, the waiter carries it back through the door, sets it on the table, and walks back for the next one. One door, back and forth. On a busy night the waiters queue up at the door — the dining hall screams for food, the kitchen screams for someone to carry plates.
Then the owner had an idea. He bricked up the small door. In its place he cut a wide pass-through in the wall and installed two lazy Susans — continuous loop turntables, one spinning from the dining hall into the kitchen, the other spinning from the kitchen out to the hall.
The rules changed.
No waiter pushes a door anymore. He writes the order on a slip, drops it into an empty slot on the inbound turntable, and the turntable carries it into the kitchen. On the other side, a kitchen runner watches the turntable, plucks each slip off, and hands it to the chef. When a dish is ready, the runner sets the plate on a slot of the outbound turntable, and it glides into the dining hall. Out front, the waiter sees the plate appear, picks it up, and delivers it.
Nobody walks through the door.
The waiter is fast now: write a slip, drop it in a slot, turn around and greet the next table — he can stuff five or six slips in a row before even glancing at the outbound turntable for finished plates. The kitchen works the same way: the chef pulls slips in batches, fires dishes in batches, never waiting for a waiter to show up one trip at a time.
The turntables are loops — a fixed number of slots, and when they come around they start over. Full means full: if the inbound turntable has no empty slots, the waiter waits until one clears. Same on the return: if the outbound turntable is packed, the kitchen waits for a waiter to pick up a plate. But as long as both sides keep moving and don’t stop, the flow never breaks.
A clever waiter discovered something else: he can pre-load slips for the signature dishes that the next table is almost certain to order — before the customer even speaks, the turntable is already carrying the order in. The food arrives so fast it feels like magic.
A new waiter once asked an old hand: “Before, when I pushed the door open, I could only carry one slip at a time, and I had to stand at the door waiting for the chef to take it. Now I can drop ten slips without leaving the dining hall. What exactly am I saving?”
The old hand smiled: “You’re saving the door. Every time you pushed it, the latch had to open, close, lock, and check your face — that whole routine took longer than cooking the dish. Now you drop the slip in a slot and you don’t even need to know where the door used to be.”
— By now you’ve probably recognized it: this is io_uring.
What it is
io_uring is an asynchronous I/O framework introduced in Linux 5.1 (2019). It shares two ring buffers between userspace and kernel: the Submission Queue (SQ) and the Completion Queue (CQ). An application fills I/O requests (read, write, accept, send, recv, …) into empty slots of the SQ. The kernel picks requests out of the SQ, executes them, and drops results into the CQ. The application then harvests results from the CQ.
The key advantage: no syscall needed per I/O operation. A traditional read/write requires one user→kernel→user context switch per call (plus security checks). Under high concurrency, this overhead — “the door” — dwarfs the actual I/O. io_uring batches many requests into a single io_uring_enter call, or in polling mode, uses zero syscalls at all; both sides communicate only through shared memory.
Why it matters
In network-intensive (Nginx, database proxies, storage engines) and disk-intensive (RocksDB, Ceph) workloads, io_uring outperforms the traditional epoll + thread-pool model by 30%–300%, because it eliminates the dominant fixed cost: syscall context switching. It is also the only general-purpose interface on Linux that delivers truly asynchronous disk I/O (the older aio subsystem only supports direct I/O with heavy restrictions). If you run high-throughput Kubernetes workloads — databases, message queues, object stores — the engine underneath is very likely already using io_uring. Understanding its ring semantics, backpressure behavior (SQ-full blocks submission), and polling-mode tradeoffs directly affects whether you tune nr_requests, iodepth, and CPU affinity correctly.
Metaphor mapping
- kitchen in the back, dining hall up front, thick wall between them → kernel space and userspace, isolated
- pushing the door, handing a slip, standing there waiting → a traditional
read/writesyscall, one context switch per call - the latch opening, closing, locking, checking your face → syscall overhead: context switch + security checks
- bricking the door, cutting a pass-through, installing two turntables →
io_uring_setup, creating the shared ring buffers - the inbound turntable (dining hall → kitchen) → the Submission Queue (SQ), where the app drops I/O requests
- the outbound turntable (kitchen → dining hall) → the Completion Queue (CQ), where the kernel drops results
- a slip dropped into a slot → writing one SQE (Submission Queue Entry)
- a plate placed on the return slot → the kernel writing one CQE (Completion Queue Entry)
- stuffing five or six slips before looking at the outbound side → batched submission, one
io_uring_enterfor many requests - turntable is a loop, full means wait → ring buffer is fixed-size, SQ-full means backpressure
- pre-loading slips for the signature dish → pre-registered buffers / fixed files, reducing per-request setup cost
- nobody walks through the door → polling mode, zero syscalls, pure shared-memory communication
- waiter and chef each working nonstop, independently → userspace and kernel running asynchronously in parallel, pipelined