Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
45 changes: 30 additions & 15 deletions docs/README.md
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
# Overlaybd Instant Container Image Pulling, No More Waiting
# Overlaybd - Instant Image Pulling, No More Waiting

<img src="assets/overlaybd_logo.svg" width="100px"/>

Expand All @@ -18,14 +18,15 @@ Overlaybd is an open-source sub-project of containerd, the industry-standard con
[graduated at CNCF](https://www.cncf.io/projects/containerd/).
The project has been integrated by many organizations world-wide, most notably
[Azure Kubernetes Service's Artifact Streaming](https://learn.microsoft.com/en-us/azure/aks/artifact-streaming-overview),
[Databricks](https://www.databricks.com/blog/booting-databricks-vms-7x-faster-serverless-compute),
[Databricks](https://www.databricks.com/blog/booting-databricks-vms-7x-faster-serverless-compute) +
[Superhuman](https://www.databricks.com/blog/how-superhuman-and-databricks-built-200k-qps-inference-platform-together),
[DeepSeek Elastic Compute (DSes)](https://arxiv.org/html/2606.19348v1),
[fly.io](https://community.fly.io/t/experimental-speedy-machine-creation-with-overlaybd/18958),
[hocus.dev](https://hocus.dev/blog/virtualizing-development-environments), etc.
[Kimi AgentEnv](https://kvcache.ai/blog/agentenv-open-sourced/),
Overlaybd can also be used for virtual machines or micro sandboxes.

<!-- Boss直聘, -->
<!-- RedNote(小红书), etc. -->

Overlaybd is the official open-sourced implementation of
two papers published in USENIX Annual Technical Conference:
Expand All @@ -43,7 +44,14 @@ Overlaybd achieves fast container startup through the following key designs:
- **On-demand (lazy) fetching**: Containers start immediately by reading image data remotely as needed,
without downloading and unpacking the entire image. A 1GB+ image can launch in under 1 second.
With proper caching / P2P transferring, overlaybd can support [launching 10,000s of containers
within seconds](articles/dadi-aliyun-2020-en.md).
within seconds](dadi-aliyun-2020-en.md).

- **Prefetching**: On-demand fetching can be combined with prefetching to warm up data
before it is needed. Two modes are supported:
- *Trace-based*: record the block access pattern of a prior run and replay it to
prefetch the same blocks ahead of the next launch.
- *File-list-based*: given a list of files required at startup, prefetch the blocks
backing those files in advance.

- **Optimized LBA lookup**: Overlaybd operates at the block level with a flat LBA address space, so
cross-layer lookup is a **single** index query regardless of layer count, whereas filesystem-based
Expand All @@ -52,7 +60,7 @@ Overlaybd achieves fast container startup through the following key designs:
Furthermore, overlaybd's multi-layer block mapping uses a highly optimized linearized B+ tree
with AVX-512 SIMD batch comparison, achieving 100s of millions of QPS on a single CPU core,
up to 10X over conventional approaches.
[Details](https://github.com/containerd/overlaybd/blob/main/docs/lsmt_lookup.md).
See [this](lsmt_lookup.md) for details.
Note that this algorithm is so fast that it is [even feasible for IP address lookup in
backbone core routers](https://www.usenix.org/conference/nsdi26/presentation/zhang-zhihao).

Expand All @@ -65,25 +73,30 @@ Overlaybd achieves fast container startup through the following key designs:
transfer is reduced enough that compressed random reads are often *faster* than
uncompressed ones — the decompression cost is less than the I/O saved.

- **High-efficiency I/O runtime**:
The service is built on [PhotonLibOS](https://github.com/alibaba/PhotonLibOS),
which is probably the fastest coroutine framework in the world.

<style>
.pic-row {
display: flex;
flex-wrap: wrap;
justify-content: flex-start;
gap: 12px 2%;
margin-left: 1.5em; /* match docsify's list padding so images align with the list text */
}
.pic {
width: 30%; margin-right: 2%;
width: 30%;
background-color: #3f3f3f;
}
</style>

- **Benchmark results**:

<div class="pic-row">
<img src="assets/cold_startup_latency.png" class="pic"/>
<img src="assets/warm_startup_latency.png" class="pic"/>
<img src="assets/startup_latency_with_prefetch.png" class="pic"/>
<img src="assets/batch_code_startup_latency.png" class="pic"/>
<img src="assets/time_pull_image.png" class="pic"/>
<img src="assets/time_launch_app.png" class="pic"/>
</div>


## Secure
Expand Down Expand Up @@ -182,16 +195,15 @@ the world:

## Open-source

Overlaybd is developed in the open under neutral governance:
Overlaybd is an open-source sub-project of [containerd](https://www.cncf.io/projects/containerd/)
(a CNCF graduated project) since 2021, released under the Apache-2.0 license.

- **Source code repositories on GitHub**:
- I/O data path: [containerd/overlaybd](https://github.com/containerd/overlaybd)
- Data I/O path: [containerd/overlaybd](https://github.com/containerd/overlaybd)
- Snapshotter & conversion tools: [containerd/accelerated-container-image](https://github.com/containerd/accelerated-container-image)
- P2P distribution: [data-accelerator/dadi-p2proxy](https://github.com/data-accelerator/dadi-p2proxy)

- **Neutral governance**: An open-source sub-project of
[containerd](https://www.cncf.io/projects/containerd/) (a CNCF graduated project) since
2021, released under the Apache-2.0 license. Originally proposed by Alibaba Cloud, it is
- **Neutral governance**: Originally proposed by Alibaba Cloud, it is
maintained in the open rather than controlled by a single vendor.

- **Active, broadening community**: Beyond the core maintainers, the project receives
Expand All @@ -203,11 +215,14 @@ Overlaybd is developed in the open under neutral governance:
implementation:
[LSMT](https://containerd.github.io/overlaybd/#/specs/lsmt.md),
[ZFile](https://containerd.github.io/overlaybd/#/specs/zfile.md).
Actually we already see several re-implementations of overlaybd / zfile, such as
[AgentEnv](https://github.com/kvcache-ai/AgentENV/tree/main/storage/overlaybd/src).

- **Join the discussion**: For synchronous communication, catch us in the `#overlaybd`
channel on the [CNCF slack](https://cloud-native.slack.com) (cloud-native.slack.com).
Everyone is welcome to join and chat.
[Get an invite to the CNCF slack.](https://communityinviter.com/apps/cloud-native/cncf)
DingTalk Group: 186405011387.

# Components

Expand Down
Binary file added docs/assets/cold_vs_warm_start.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/assets/p2p_distribution.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/assets/startup_flow_comparison.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/assets/startup_latency_10k.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Binary file added docs/assets/stress_test_10s.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
53 changes: 53 additions & 0 deletions docs/dadi-aliyun-2020-en.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
# Launching 10,000 Containers in Seconds: Inside Alibaba Cloud's Container Image Acceleration

> Author: Mu Huan | Published: 2020-01-08
> Source: https://developer.aliyun.com/article/742103

## Introduction

Alibaba Cloud's container and storage teams collaborated to leverage the **DADI accelerator**, enabling on-demand image reading and P2P distribution. The result: **10,000 containers launched in 3.01 seconds**, completely eliminating the minutes-long cold start wait and the network congestion caused by massive parallel image distribution from registries.

## Background and Pain Points

When handling traffic spikes from events like shopping festivals, flash sales, and live streaming, containers must scale up within an extremely short time window.

- **Warm start**: With a local image cache, instantiation is very fast.
- **Cold start pain**: Without a local image, the full image must be downloaded and decompressed from the registry (constrained by network and disk performance), taking minutes. Large-scale batch cold starts can even render the registry unresponsive due to network congestion.
- **Data redundancy**: A full container image is often hundreds of MB or even GBs, yet the data actually needed for startup may be a tiny fraction (e.g., a test image of 894MB required only 15MB to start — about 1.6%).

## Test Scenario and Results

Tests were conducted on Alibaba Cloud Container Service for Kubernetes (ACK), using a cluster of **1,000 nodes with 4 cores and 8GB RAM each**.

- **10,000-container launch**: Starting 10,000 containers took only **3.01 seconds** (p999 latency: 2.97s).

![Startup latency for 10,000 containers](assets/startup_latency_10k.png)

- **Cold vs. warm start**: Warm start took 2.91s (p999: 2.56s). DADI cold start, fetching data on-demand from the P2P network, reduced disk I/O pressure and produced fewer long-tail containers.

![Cold vs. warm start latency comparison](assets/cold_vs_warm_start.png)

- **Time-limited stress test**: Within a 10-second window, **59,997 containers** were launched across 1,000 hosts (the 60,000th container started at 10.06s).

![10-second stress test](assets/stress_test_10s.png)

## Core Technology: The DADI Accelerator

The DADI accelerator bypasses the traditional "download image → decompress image → start container" pipeline. Its core principle is "on-demand reading," achieved through three key optimizations:

![Container startup flow comparison](assets/startup_flow_comparison.png)

1. **Image format optimization**: A new image format with a built-in index allows direct access to required data without downloading or decompressing the entire image.
2. **On-demand P2P data fetching**: A tree-structured P2P network distributes data. A small number of P2P root nodes fetch from the registry, while other host nodes transfer data among themselves, greatly reducing single-point registry load and network congestion.

![DADI P2P data distribution](assets/p2p_distribution.png)

3. **Efficient decompression**: A new compression file format enables on-demand decompression of only the data actually accessed, with decompression overhead that is nearly negligible.

**How it works**:

> When launching a container with DADI, only a few KB of image metadata is downloaded from the registry. A virtual device — the **Overlay Block Device** — is created and mounted to the container's working directory. The Docker engine considers the image fully loaded. The actual image data needed at runtime is then fetched on-demand from the local cache or upstream nodes in the P2P network.

## Conclusion

The deep integration of Alibaba Cloud ACK and DADI achieved launching tens of thousands of containers in seconds, gracefully handling large-scale application scaling and deployments. This technology will also serve as a key enabler for **serverless container** startup acceleration in the future.
53 changes: 53 additions & 0 deletions docs/dadi-aliyun-2020.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
# 秒级启动万个容器,探秘阿里云容器镜像加速黑科技

> 作者:木环 | 发布日期:2020-01-08
> 来源:https://developer.aliyun.com/article/742103

## 简介

阿里云容器与存储团队展开合作,利用 **DADI加速器** 支持镜像按需读取和P2P分发,实现 **3.01秒启动10000个容器**,完美杜绝容器冷启动的数分钟漫长等待,以及镜像仓库大规模并行分发场景下的网络拥堵。

## 背景与痛点

在应对年货节、秒杀、直播等流量瞬间暴增的业务时,必须在极短时间内扩容容器。

- **热启动**:本地有镜像缓存,实例化时间很短。
- **冷启动痛点**:本地无镜像时,需从Registry下载并解压完整镜像(受网络和磁盘性能影响),耗时数分钟;大规模批量冷启动甚至会导致Registry因网络拥堵而无法响应。
- **数据冗余**:完整容器镜像往往高达数百M甚至上G,但启动所需数据可能仅占极小比例(例如测试镜像894M,启动仅需15M,占比约1.6%)。

## 测试场景与数据表现

测试在阿里云容器服务ACK中进行,场景为 **1000台4核8G节点** 组成的Kubernetes集群。

- **万级容器启动**:启动10000个容器仅需 **3.01秒**(p999分位耗时2.97秒)。

![1万个容器的启动耗时](assets/startup_latency_10k.png)

- **冷热启动对比**:热启动耗时2.91秒(p999耗时2.56秒)。DADI冷启动由于数据按需从P2P网络获取,减轻了磁盘IO压力,长尾容器较少。

![冷热启动耗时对比](assets/cold_vs_warm_start.png)

- **限时摸高测试**:10秒限制内,利用1000台宿主机启动了 **59997个容器**(10.06秒时第6万容器启动完毕)。

![限时10秒摸高测试](assets/stress_test_10s.png)

## 核心黑科技:DADI 加速器

DADI加速器避免了传统"下载镜像 -> 解压镜像 -> 启动容器"的步骤,其核心在于"按需读取",主要包含以下三点优化:

![容器启动流程对比](assets/startup_flow_comparison.png)

1. **镜像格式优化**:设计了一种包含索引的新型镜像格式,无需下载和解压完整镜像即可直接访问所需数据。
2. **按需P2P数据读取**:利用树形P2P网络进行数据分发。少数P2P根节点从Registry获取数据,其他宿主机节点之间相互传输,极大减轻了单点Registry的负载和网络拥堵。

![DADI P2P数据分发](assets/p2p_distribution.png)

3. **高效的解压缩算法**:提供新型压缩文件格式,可按需单独解压用户实际访问的数据,解压时间开销几乎可忽略不计。

**运行机制**:

> 利用DADI方案启动容器时,仅从镜像Registry下载几KB的镜像元数据,并创建虚拟设备 **Overlay Block Device** 挂载到容器工作目录上,Docker引擎会认为镜像已经加载完毕。在容器启动时所需的镜像数据则从本地缓存或者P2P网络的上游节点按需下载。

## 总结

此次阿里云ACK和DADI的深度整合,实现了秒级启动万个容器,从容应对大规模应用扩容和发布。相关技术在未来也将成为 **Serverless容器** 的启动加速利器。
3 changes: 2 additions & 1 deletion docs/index.html
Original file line number Diff line number Diff line change
Expand Up @@ -7,8 +7,9 @@
<meta name="description" content="Description">
<meta name="viewport" content="width=device-width, initial-scale=1.0, minimum-scale=1.0">
<link rel="icon" type="image/svg+xml" href="assets/overlaybd_logo.svg">
<!-- <link rel="stylesheet" href="https://cdn.jsdelivr.net/npm/docsify-themeable@0/dist/css/theme-simple.css"> -->
<link rel="stylesheet" href="//cdn.jsdelivr.net/npm/docsify@5/dist/themes/core.min.css" />
<link rel="stylesheet" href="//cdn.jsdelivr.net/npm/docsify@5/dist/themes/addons/core-dark.min.css" />
<!-- <link rel="stylesheet" href="//cdn.jsdelivr.net/npm/docsify@5/dist/themes/addons/buble.min.css" /> -->
<style>
.markdown-section tr:nth-child(2n) {
background-color: transparent;
Expand Down
Loading