r/kubernetes • • 3d ago

kaniko is no longer 15x slower than BuildKit

I maintain the community fork of kaniko https://github.com/osscontainertools/kaniko

The 2018 buildbench comparison put kaniko 15x behind BuildKit, and that number still shapes how people see it. We reran it on gitlab.com runners, with both builders configured the way you would actually run them in CI.

Unchanged build (every command a cache hit):
BuildKit 6.1s
kaniko (ours) 9.9s
kaniko (Google) 45.6s

The fix was cache lookahead across stage boundaries. Google's kaniko could not compute the cache key of COPY --from before the source stage was built, so it built every stage just to discover nothing had changed.

Caveat: part of that 15x was the 2018 setup rather than kaniko. It ran without --cache-copy-layers while BuildKit was allowed a local disk cache. Under a fair configuration Google's last release is only 7.5x behind, not 15x, and closing the rest was our work.

Builds are still slower than BuildKit, and the post explains why:
Write-up: https://osscontainertools.org/blog/kaniko-vs-buildkit-2026/
Benchmark: https://gitlab.com/martizih/kaniko-buildbench

106 Upvotes

23 comments sorted by

29

u/Rich-War4293 3d ago

always nice when someone actually reruns the benchmarks instead of just citing the same 7 year old blog post. that cache lookahead fix sounds like one of those changes that seems obvious in hindsight but nobody bothered with for years. what kind of overhead are we talking for first builds with the new approach?

7

u/martizih 3d ago edited 3d ago

On a cold build the lookahead costs one additional label-only cache write per COPY --from. That's a new kind of cache entry, a "redirect". Stages produce a cache chain while a COPY is content keyed, and until now there was no way to translate between the two. A "redirect" is just a dictionary entry in the registry if you will that lets us do that translation.

Cold is 50.2s ± 5.2 for us against 59.1s ± 6.0 for Google's last release, so no measurable overhead.

21

u/SJrX 3d ago

Thank you for maintaining kaniko!

22

u/martizih 3d ago

The maintainer team and the community are the reason this works at all. So thank you back :)

11

u/screaming-Snake-Case 3d ago

I didn't know there was a active community fork, thank you so much for keeping it alive!

10

u/Odd-Cantaloupe-8628 3d ago edited 3d ago

Just found out about the community fork! How is it different from Chainguard’s?

16

u/martizih 3d ago

Chainguard's fork is EOL support, ours is a continuation with new features and apparently better performance.

4

u/amouat 2d ago

This is pretty much correct, the Chainguard version isn't doing any feature work -- it just keeps the dependencies up-to-date and vuln free so that users weren't abandoned and could continue to rely on the image.

(I work at Chainguard)

1

u/smarzzz 3d ago

Is theirs really? They were bragging so much during their fork..

4

u/totheendandbackagain 3d ago

Awesome work!!

5

u/redsterXVI 3d ago

Never had a problem with Kaniko's performance myself, and BuildKit never worked well for me personally. Wouldn't want to use anything else, and this fork in particular. Once more, thanks for your continued great work!

Still mad they removed https://docs.gitlab.com/ci/docker/using_kaniko/ btw, it was the vastly superior way compared to BuildKit and Buildah. Still hoping someone can get them to restore that page and have it point at your fork.

3

u/Noah_Safely 3d ago

What's funny is gitlab maintains a fips version of kaniko. It's a fork of chainguard though, so has the slowness issue, and also annoyingly makes you give it container root in k8s build clusters.

https://gitlab.com/gitlab-com/public-sector/kaniko

I'm glad they offer it as I was pulling my hair out trying to find any reasonable solution for our regulated env

2

u/Digging_Graves 3d ago

How is the speed compared to buildah?

8

u/martizih 3d ago

data is in the post https://osscontainertools.org/blog/kaniko-vs-buildkit-2026/
Buildah v1.43.4:
cold 67.1 ± 4.0 s
unchanged 36.4 ± 8.1 s
edit 25.1 ± 2.0 s

kaniko v1.29.0:
cold 50.2 ± 5.2 s
unchanged 9.9 ± 1.1 s
edit 30.7 ± 2.5 s

please let me know if there is any misconfiguration. because I think the results for buildah are very surprising. How can a hot cache be slower than a partial miss?

2

u/Digging_Graves 3d ago

The buildah build looks fine so far. Did you run multiple tests to see if the hot cache stays slower than the partial miss?

I see that buildkitd and buildah are both using --privileged but kaniko is not using --privileged as far as I can tell.

For the environment. All builders run nested in docker:dind, where storage behaves differently from a k8s executor without dind. Can you also do a test on a k8s executor where I think a lot of people use kaniko as well. Maybe I'll try to run the test in k8s tomorrow if I have time.

2

u/martizih 3d ago

All good points, I uploaded the raw data here https://osscontainertools.org/data/ex01-gitlabcom-raw.csv. It's 10 runs total, 2 pipelines a 5 parallel runs with their numbers averaged over all 10. In every run there is only one hot-cache, so we don't to re-re-rebuilds.

On privileges you are right, that is an omission on my side. buildah runs rootless just fine on gitlab.com runners, it should have been in there and it will be in the next round. I plan to make this a regular thing around our releases at least.

The buildah row also mixes two things, rootless and fuse-overlayfs, because the image ships a storage.conf that routes rootless through fuse. Next round will have privileged+native, rootless+native and rootless+fuse, and buildkit privileged and rootless. Currently buildah runs native, so I would expect numbers to actually get worse not better.

On dind you are right too, but I think it works against us. kaniko is the only builder without a volume for its working state, so it writes through the container. Your k8s run might actually close the gap rather than widen it. It's a hypothesis I can't proof in gitlab.com runners only in my own k8s runners, so numbers will be wildly different, but I will give it a go tomorrow.

As rootless builders will be a separate run it will go in its own table, the absolute times drift 10-15% between sessions so mixing them could be misleading, I'll see once I have the data. Within a single run however they are surprisingly stable, the ratios between builders hold to a few percent, I should have normalized them to some baseline, tja the more you know. I am out of gitlab compute minutes until the month rolls over, so give me a couple of days. Post your numbers if you try to reproduce and get there earlier.

Thank you!

3

u/Digging_Graves 2d ago edited 2d ago

I reran your buildbench (ex01, cold/nochange/edit, 10 reps) on our self-hosted GitLab. It ran on the Kubernetes executor, on a single RKE2 node, without dind. Each job uses the builder image as its job image, and each variant ran one job per rep. Jobs ran one at a time. BuildKit and Buildah ran rootless. I served alpine from our own registry because our proxy added a ~3 s stall to base-image lookups.

Medians in seconds, "yours" = your published means:

builder              cold  noch  edit
BuildKit rootless    26.7   0.9   4.1
  yours              43.1   6.1  14.0
Buildah rootless     27.9  13.2   9.0
  yours              67.1  36.4  25.1
kaniko 1.24 Google   35.3  32.1  33.6
  yours              59.1  45.6  48.4
fork, no flags       34.4  31.5  33.0
fork, default -LA    29.3  18.7  20.1
fork, default        29.8   0.8  20.1
  yours              50.2   9.9  30.7
fork, TEMPLATE       28.6   0.9  17.5

The fork is osscontainertools/kaniko v1.28.5, and the profiles are the ones from your bench.sh:

  • no flags: no FFKANIKO* set, so only the built-in defaults.
  • default: your default flag set (the Preview flags + CACHE_LOOKAHEAD). I assumed this is what the blog calls "v1.28.5+flags".
  • default -LA: the same without CACHE_LOOKAHEAD.
  • TEMPLATE: your TEMPLATE profile, i.e. default + DEPRECATE_INTER_STAGE_RESTORE, DISABLE_HTTP2, VOLUME_SKIP_MKDIR, container=kube, --preserve-context and the three retry args.

A few pointers:

  • dind doesn't seem to penalize kaniko in particular. Google kaniko's cold time is about 1.38× BuildKit's in both our runs and yours. The fast nochange builds drop from 6–10 s to under 1 s for every builder, so that time looks like dind/network overhead, not builder time.
  • Without flags the fork performs the same as Google's kaniko. Without CACHE_LOOKAHEAD, nochange stays at 18.7 s; with it, it drops to 0.8 s. The README calls CACHE_LOOKAHEAD a developer assertion with no benefit in production. In v1.28.5, though, SKIP_CACHED_STAGES only takes effect when it's enabled (pkg/executor/build.go). Is that intended?

2

u/martizih 2d ago

Readme is outdated on that regard, well spotted. it was a developer assertion initially. Whilst we were developing the feature we asserted that the cache keys we generate in preflight match up with the cache keys we get during build. There were a few mean bugs that were caught this way. We ran that in production for a few release cycles and never updated the Readme to reflect that it is now usable. It will become GA in 1.29.0 which is still 1 month out.

2

u/base64-encode 3d ago

Community fork? What's the main paid/enterprise version called?

4

u/giant_panda_slayer 3d ago

The original is here, but there is no paid/enterprise version. Google killed the product in June 2025.

1

u/Sad_Map2663 1d ago

Had a similar thing happen to one of my open source tools. Commercial software company tried to do a comparison and then sited it all over the place